📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Data Science Full Course 2025 | Data Science Tutorial | Data Science Training Course | Simplilearn

Simplilearn11:00:55

Transcription

Hey everyone, welcome to Simply Learn data science full course. Did you know data science helps you uncover hidden patterns and solve real-world problems?

So in today's course, you will learn how to turn raw data into valuable insights. Think of data as a treasure chest, and data science is the key that unlocks its value. This course blends programming, statistics, and practical tools to help you transform numbers into smart decisions. We'll dive into the fundamentals of data science, explore probability and statistics, and also introduce advanced tools like generative AI. We'll also gain hands-on projects to give you the experience to tackle real-world challenges.

But before we begin, if you're looking to fast-track your career in data science, then the professional certificate course in data science is your perfect opportunity. In collaboration with the ENICT Academy IT Kpur, it offers a world-class program that combines expert guidance and practical learning. You'll also earn a prestigious program certificate from IT Kur and master classes delivered by the distinguished faculty. This course covers the latest tools and technologies like charge GPT, Python, PowerBI, Tableau, ensuring you stay ahead in this field. You can find the course link in the description box below and in the pin comment. So hurry up and enroll now.

Are you one of the many who dreams of becoming a data scientist? Keep watching this video if you're passionate about data science because we will tell you how does it really work under the hood. Emma is a data scientist. Let's see how a day in her life goes while she's working on a data science project.

Well, it is very important to understand the business problem first. In her meeting with the clients, Emma asks relevant questions, understands and defines objectives for the problem that needs to be tackled. She's a curious soul who asks a lot of advice—one of the many traits of a good data scientist.

Now she gears up for data acquisition. To gather and scrape data from multiple sources like web servers, logs, databases, APIs, and online repositories. Oh, it seems like finding the right data takes both time and effort.

After the data is gathered comes data preparation. This step involves data cleaning and data transformation. Data cleaning is the most time-consuming process as it involves handling many complex scenarios. Here Emma deals with inconsistent data types, misspelled attributes, missing values, duplicate values, and whatnot. Then in data transformation, she modifies the data based on defined mapping rules. In a project, ETL tools like talent and Informatica are used to perform complex transformations that help the team to understand the data structure better.

Then understanding what you actually can do with your data is very crucial. For that, Emma does exploratory data analysis. With the help of EDA, she defines and refines the selection of feature variables that will be used in the model development. But what if Emma skips this step? She might end up choosing the wrong variables, which will produce an inaccurate model. Thus, exploratory data analysis becomes the most important step.

Now she proceeds to the core activity of a data science project, which is data modeling. She repetitively applies diverse machine learning techniques like KN&N decision tree knives base to the data to identify the model that best fits the business requirements. She trains the models on the training data set and tests them to select the best performing model. Emma prefers Python for modeling the data. However, it can also be done using R and SAS.

Well, the trickiest part is not yet over: visualization and communication. Emma meets the clients again to communicate the business findings in a simple and effective manner to convince the stakeholders. She uses tools like Tableau, PowerBI, and ClickView that can help her in creating powerful reports and dashboards.

And then finally, she deploys and maintains the model. She tests the selected model in a pre-production environment before deploying it in the production environment, which is the best practice, right? After successfully deploying it, she uses reports and dashboards to get real-time analytics. Further, she also monitors and maintains the project's performance.

Well, that's how Emma completes the data science project. We have seen the daily routine of a data scientist is a whole lot of fun, has a lot of interesting aspects, and comes with its own share of challenges.

Now let's see how data science is changing the world. Data science techniques along with genomic data provide a deeper understanding of genetic issues in reaction to particular drugs and diseases. Logistic companies like DHL, FedEx have discovered the best routes to ship, the best-suited time to deliver, the best mode of transport to choose, thus leading to cost efficiency. With data science, it is possible to not only predict employee attrition but also to understand the key variables that influence employee turnover. Also, the airline companies can now easily predict flight delay and notify the passengers beforehand to enhance their travel experience.

Well, if you're wondering, there are various roles offered to a data scientist like data analyst, machine learning engineer, deep learning engineer, data engineer, and of course, data scientist. The median base salaries of a data scientist can range from $95,000 to $165,000.

So that was about the data science. Are you ready to be a data scientist? If yes, then start today. The world of data needs you.

Have you ever wondered how your favorite online store seems to know exactly what you are looking for? Every time you browse, add to cart or wish list an item, you are leaving clues about your style, favorite colors, brands, and even shopping times. Data scientists jump in, analyze these patterns, and create a super personalized shopping experience. Suddenly, the store is showing you just the right pieces at just the right time. Almost like it's reading your mind. That's data science—turning your clicks into a shopping spree crafted just for you.

Hello everyone. Welcome back to Simply Learn's YouTube channel. If you're already a data science enthusiast or just got curious about this exciting field, you're in the right place. Today in this video, I'm diving into 10 essential steps to help you become the next in-demand data scientist and land that dream job. No more waiting. Let's dive right in and get you on the path to your future in data science.

So, let's see the 10 essential steps to become the next data scientist in demand. Step number one is programming languages. Starting with Python is a beginner is a great move because it's simple, versatile, and widely used in data science. Python's straightforward syntax makes it beginner-friendly, helping you grasp programming basics quickly and dive into data science libraries like pandas, numpy, and mattplotive with ease. Adding R to your skill set is valuable because it excels at statistical analysis and data visualization, two essential parts of data science. You can be comfortable with Python and R within a month or two.

So moving on to the next step that is version control system. Learning a version control system like Git is essential because it allows you to track, manage, and collaborate and code effectively. With Git, you can save different versions of your work, making it easy to backtrack if something goes wrong or to experiment without losing progress. This is especially useful when working with complex data science projects where you might try out different models of analysis techniques. One or two weeks of practice along with Python and R is good to get started.

Now moving on to the third step that is data structures and algorithms. Learning data structures and algorithms is crucial for becoming a data scientist because they provide the foundation for efficient data handling and problem-solving. Data structures like arrays, stacks, queues, and trees help you store and organize data in ways that make it easier and faster to access, process, and analyze. Algorithms, on the other hand, give you strategies to perform tasks like searching, sorting, and optimizing data operations which are essential for handling large data sets. While many candidates struggle with the essay, mastering it gives you an edge, helping you stand out in the interviews and shine as a skilled data scientist capable of tackling the toughest data problems. Spend about two months in this, you will get in the shape for sure.

Now moving on to the step number four that is SQL. Learning SQL is essential for data scientists because it enables you to access, manage, and manipulate data directly within databases where most real-world data resides. With SQL, you can create new tables, alter existing ones, delete unnecessary records, and run queries to filter, sort, and aggregate data. These abilities allow you to retrieve, clean, and organize data effectively—core skills needed for any data science role. It's easy, and you don't have to spend more than a month to have a deep understanding of it.

Now moving on to the fifth step that is mathematics and statistics. Mathematics and statistics are essential for data science because they form the backbone of data analysis, model building, and interpretation. Topics like linear algebra, calculus, probability, and statistics give data scientists the tools to understand data patterns, perform accurate analysis, and make data-driven decisions. Mastering these areas enables you to build robust models, validate results, and tackle complex problems confidently, making you a well-rounded and skilled data scientist. Make sure you spend two months to grasp these topics.

Now moving on to the step number six that is data pre-processing and visualization. Learning data pre-processing and visualization is essential for a data scientist because these skills make your data accurate, insightful, and easy to understand. Python libraries like NumPy and Panders are crucial for manipulating and creating data, enabling you to handle missing values, filter out noise, and prepare data for analysis. Once the data is ready, visualization lets you uncover patterns and communicate results effectively. Libraries like Mattplot tip and Seaborn help create clear, impactful visuals, allowing you to interpret trends and convey insights in a way that's easily understood by others. Together, these tools make data pre-processing and visualization fundamentals for effective data science. If you have a solid foundation on Python and mathematics, you will get a good understanding of data pre-processing and visualization in a month or two.

Now moving on to the seventh step that is machine learning fundamentals. Machine learning fundamentals involve understanding how algorithms enable computers to learn from data and make predictions or decisions without explicit programming. The two main categories are supervised learning and unsupervised learning. In supervised learning, models are trained on labelled data to make predictions while in unsupervised learning models find patterns in unlabelled data. Popular tools like TensorFlow, PyTorch help build and train complex models, especially for deep learning. While scikit-learn is essentially used for simpler machine learning algorithms and data pre-processing. These tools make it easier to implement machine learning fundamentals effectively and build intelligent data-driven decisions. Dedicate about three months to understand the core of machine learning.

Now coming to the next step that is deep learning. Deep learning is a subset of machine learning that focuses on algorithms inspired by the structures of the human brain called neural networks. Deep learning uses neural networks with multiple layers—often dozens or hundreds—to learn complex patterns from large data sets. Specialized types like convolutional neural networks, that is CNN's, are great for image processing while recurrent neural networks, RNNs, are used for sequence data like text or time series. Essential tools like TensorFlow, PyTorch make building, training, and deploying deep learning models more accessible, allowing you to create powerful AI solutions across various domains. I think it will take about 2 months to have a good hold on deep learning concepts and how to implement them.

Now moving on to the ninth step that is specializations. Once you have grasped the deep learning, it's like reaching a new level as a data scientist. Just as doctors specialize in areas like nephrology and cardiology, data scientists often choose to specialize in fields like natural language processing or computer vision. Natural language processing focuses on teaching machines to understand and generate human language, enabling applications like chatbots, sentiment analysis, and language translation. It's about making computers read, write, and even interpret human emotions through text or speech. Computer vision, on the other hand, is all about enabling machines to see and interpret images or videos. This field powers innovations like facial recognition, object detection, and autonomous driving. Now you don't need to learn both. You can choose what interests you the most. Now spend one to two months diving deep into one of these areas.

Now moving on to the last but not the least step that is big data. Big data refers to extremely large volumes of data generated rapidly from sources like social media and sensors. For data scientists, learning to handle big data is crucial as it requires specialized tools like Hadoop and Spark to analyze and extract insights effectively. With companies relying on data-driven decisions, big data skills make you a highly in-demand professional in the field. Focus for about 2 months and you will be able to spot trends and patterns from data sets very easily.

Once you're ready, it's time to build a killer resume packed with projects that showcase your new skills. Start applying to jobs on platforms like Noy and Indate and supercharge your LinkedIn. Connect with data scientists. See what skills they are mastering and learn from their journeys as well. Keep sharpening your own skills, and when the time comes, you will be ready to crush those interviews and land your dream data scientist role in 2025.

[Music]

Learning objectives. Welcome to math refresher, probability, and statistics. In this lesson, we are going to explain the concepts of statistics and probability. Describe conditional probability. Define the chain rule of probability. Discuss the measure of variance. Identify the types of Gaussian distribution.

Basic of statistics and probability. Probability and statistics. Data science relies heavily on estimates and predictions. A significant portion of data science is made up of evaluations and forecasts. Statistical methods are used to make estimates for further analysis. Probability theory is helpful for making predictions. Statistical methods are highly dependent on probability theory. And all probability and statistics are dependent on data. Data is information acquired for reference or research via observations, facts, and measurements. Data is a set of facts structured in the form that computers can interpret, such as numbers, words, estimations, and views.

Importance of data. Data aids in seeing more about the information by identifying possible connections between two features. Data assists in the detection of distortion by uncovering hidden patterns based on prior information patterns. Data may be utilized to anticipate the future or predict the current state of affairs. Also, data aids in determining whether two pieces of information have any instance in common or not.

Types of data. Data might be quantitative—that is, data that can be measured or counted in numbers—or it may be qualitative, which is data which is generally divided into groups or, in simpler words, which cannot be counted or measured in numbers. Let's consider an example. A customer information data of a bank may contain quantitative and qualitative data. Consider this snapshot where we have customer ID, surname, geography, gender, age, balance, has C or card, is active member. Amongst these variables, we can see surname is mostly qualitative as it cannot be counted and measured in numbers. Geography and gender are also qualitative as they cannot be counted in numbers and are mostly groups. has C or card, that is has credit card, and is active member although are containing numerical in form but these are categorical; that means these have been divided into groups of one and zero that represent yes and no as an answer; hence these two variables are also qualitative. Customer ID is again, although a numerical data, however the significance or intuition behind Customer ID is categorical. Hence, it may be kept in the qualitative data also. However, age and balance—these are numerical information which have been measured or counted, and numerical operations can be performed on them. Hence, these are under quantitative data categories.

Introduction to descriptive statistics. Descriptive statistics. A descriptive measurement is a summary measure that quantitatively portrays the most important features of a set of data, allowing for a better comprehension of the information. Data can be measured as different levels. The levels of measurement describe the nature of information stored in the data assigned to the variables. Qualitative data can be measured as nominal or ordinal. Quantitative data can be measured in terms of interval and ratio type.

Nominal data. The data is categorized using names, labels, or qualities. For example, brand name, zip code, and gender.

Ordinal data can be arranged in order or ranked and can be compared. Examples include grades, star reviews, position, and race, and date.

Interval data is the data that is ordered and has meaningful differences between the data points. Example, temperature in Celsius and year of birth.

Ratio data is similar to the interval level with the added property of inherent zero. Mathematical calculations can be performed on both interval as well as ratio data. For example, height, age, and weight.

Population versus sample. Before analyzing the data, it's important to figure out if it's from a population or a sample. Population is a collection of all available items as well as each unit in our study. Sample is a subset of the population that contains only a few units of the population. Population data is used for study when the data pool is very small and can give all the required information. Samples are collected randomly and represent the entire population in the best possible way.

Measures of central tendency. The central tendency is a single value that aids in the description of the data by determining its center position. Measures of central tendency are sometimes known as summary statistics or measures of central location. The most popular measurements of central tendency are mean, median, and mode. The normal distribution is a bell-shaped symmetrical distribution in which mean, median, and mode all are equal. The curve over here shows the bell-shaped curve or the normal distribution of variable X. The point over here, that is X1, is the point which represents the mean, median, and mode of this distribution.

Mean. Mean is calculated by dividing the sum of all data values by the total number of data values. It gets affected when there are unusual or extreme values. It is sensitive to the outliers. Mean can be calculated as summation over all the values of X in a collection divided by the size of the collection. For example, we have a collection where we have values as 7, 3, 4, 1, 6, and 7. We find out the sum of these values, which is 28, and there are a total of six values. So 28 / 6 gives us a mean value of 4.66.

Median. It is the middle value in the set of the data that has been sorted in ascending order. It is a better alternative to mean since it is less impacted by outliers and skewness. It is closer to the actual central value. Median is calculated differently for different sizes of data, differentiated as if the total number of values is odd or if the total number of values is even. If the size of the data is odd. For example, in this case, we have five elements. After sorting, whatever middle value we get, that means n + 1 by 2 term; in this case, 5 + 1 / 2, that is the third term, which is four, is the median value. In case when the total number of values is even, like here there are six values, the average or the mean of the two central values is considered as the median. In this case, the median is the mean of six and four, which is five.

Mode. Mode represents the most common value in the data set. It is not at all affected by extreme observations. It is the best measure of central tendency for highly skewed or non-normal distribution. Mode for categorical data is determined by estimating the frequencies for each categories, and then the category with the highest frequency is considered to be mode. Like in this case, seven has the highest frequency. Hence, seven becomes the mode value. However, in case of continuous data or quantitative data, the calculation of mode is slightly different. The first step in calculation of mode is dividing the data into classes which are equal, with then getting the frequency of data points lying in within that range of classes, and finally selecting the class with the highest frequency. Using the range of that class and the frequencies, we can get the final mode value. Using the formula L plus FM minus F_1 multiplied to H / F minus F_1 plus FM minus F_2. Here L is the lower limit or the lower observation of the mode class. H is the size of the mode class. FM is the frequency of the mode class. F_1 is the frequency of the class preceding to mode, and F_2 is the frequency of the class succeeding to mode. This gives us the final mode value.

Mean versus expectation. Now let's talk about mean versus expectation. So in general, we use the expected value or expectation when we want to calculate the mean of a probability distribution that represents the average value we expect to occur before collecting any data. And mean, on the other hand, mean is basically used when we want to calculate the average value of a given sample. This represents the average value of raw data that we may have already collected. We can understand this by using a simple example. Now to calculate the expected value of this probability distribution, we can use a specific formula from the previous discussion. This is going to be the expected value where X is going to be the data value and this PX is the probability of value. For example, we could calculate the expected value for this probability distribution to be as shown. So here it will be 1.45 goals. So this represents the expected number of goals that the team will score in any given game. And then if you talk about calculating mean, so we typically calculate the mean after we have

Actually, we collected raw data. For example, suppose we record the number of goals that a soccer team will score in 15 different games.

Now, to calculate the mean number of goals scored per game, we can use the following formula: where sum of x is basically the sum of all the goals divided by n, and n is the number of records, or we can say the sample size. It is as shown on the screen. So this represents the mean number of goals scored per game by the team.

Measures of asymmetry. The difference between the three distinct curves can be studied in this image. The central curve is the normal or no skewness curve. Here, mean, median, and mode all lie on the same point. This normal curve is symmetrical about its mean, median, and mode. That means the left-hand side of the curve is a mirror image of the right-hand side of the curve. However, in case of negatively skewed data, the tail is elongated on the left-hand side, and the mean is smaller than the mode and the median values, or is on the left-hand side of the mode. Hence, indicating that the outliers are in the negative direction. On the other hand, in case of positively skewed data, the data is concentrated on the left-hand side of the curve. While the tail is elongated or longer on the right-hand side of the curve, the mean is greater than the mode and median, or is on the right-hand side of the mode and median, indicating that the outliers are in the positive direction.

Let's consider an example. The graph here shows the global income distribution for the year 2003, 2013, and a projection for 2035. If we see the global income distribution statistics for 2003, it is highly right-skewed. We can observe in the previous graph that in 2003 the mean of $3,451 was higher than the median of $1,090. The global income is definitely not evenly distributed. The majority of people make less than $2,000 each year, while only a small percentage of the population earns more than $14,000.

Measures of variability. Measures of variability. Dispersion. The measure of central tendencies provide a single value that addresses the full worth. However, the central tendency cannot depict the viewpoint entirely. The metric of dispersion helps us focus on the inconsistency in the data spread. Measures of dispersion describe the spread of the data. The range, interquartile range, standard deviation, and variance are examples of dispersion measures.

Range. The range of distribution is the difference between the largest and the smallest amount of data. The range, for example, does not include all of a series' positive aspects. It concentrates on the most shocking aspects and ignores those that aren't considered critical. For example, for a set 13, 33, 45, 67, 70. The range is 57. That is the maximum of this, which is 70 minus the minimum over here, which is 13.

Variance. Variance is the average of all squared deviations. It is defined as the sum of squared distance between each point and the mean, or the dispersion around the mean. The standard deviation is used as variance suffers from a unit difference. Variance can be computed as sigma square summation over x - mu^2 divided by n, where mu is the mean of the data, x is the individual data point, and n is the size of the data. This representation is for a population data; for a sample data, variance can be computed as (X - Xbar)^2 summation over it divided by n - 1. Here, Xbar is the mean of these sample data, and n is the sample size. The units of values and variance are not equal. So another variability measure is used.

Standard deviation. Standard deviation is a statistical term used to measure the amount of variability or dispersion around a mean. The standard deviation is calculated as the square root of variance. It depicts the concentration of the data around the mean of the data set. Standard deviation, as indicated previously, can be computed as the square root of variance for a population data. Standard deviation sigma can be computed as the square root of summation over (xi - mu)^2 / n, where mu is the mean of the data, xi are the data points, and n is the size.

Let's consider an example. Let's find out the mean, variance, and standard deviation for this data. The data values are 3, 5, 6, 9, and 10. To find out the mean, we first find the sum of all these data values, that is 33, and divide it by the count, which is five. We get the mean of 6.6. To compute the variance, we start by computing the deviation. That is X minus the mean of X. Here, 3 is one of the values of the data, and 6.6 is the mean. So (3 - 6.6)^2, and we do that to find out the sum of all the deviations divided by the count, which is five. We end up getting an overall variance of 6.64. Standard deviation, as we know, is measured as the square root of variance, that is the square root of 6.64, which amounts to 2.576.

Measures of relationship. Measures of relationship covariance. Covariance is the measure of joint variability of two variables. It measures the direction of the relationship between the variables. It determines if one variable will cause the other to alter in the same way. Covariance between variable X and Y can be computed as summation over the product of (Xi - Xbar) and (Yi - Ybar) the whole divided by N - 1. Here, Xbar and Ybar are the mean of X and Y respectively. The value of covariance can range from minus infinity to plus infinity.

Correlation. Correlation is normalized covariance. It measures the strength of association between two variables. The most common measure for correlation is the Pearson correlation coefficient. Correlation between two variables X and Y can be measured with respect to covariance as covariance between X and Y divided by the standard deviation of X and standard deviation of Y. The value of correlation ranges from -1 to +1.

Types of correlation. Correlation can be either a positive correlation, zero correlation, or a negative correlation. The first picture over here represents a perfect positive correlation, wherein a straight line with a positive slope is representing the relationship between the two variables. Zero correlation means that the line representing the relationship between the two variables is horizontal to the x-axis. Perfect negative correlation can be represented by a straight line with a negative slope. Correlation equals to 1 implies a positive relationship; that is, when one variable increases, the other variable also increases. A correlation value of -1 implies a negative relationship; that is, when one variable increases, the other decreases. The correlation coefficient of zero shows that the variables are completely independent of each other.

Let's consider an example. Here, we have two variables: height and weight. To compute the correlation between height and weight, we use the correlation formula as covariance of X and Y divided by standard deviation of X and standard deviation of Y. Here, height is the X variable, and weight is the Y variable. First, to compute covariance, we compute the (x - xbar) and (y - ybar) values and then the product of them. We then compute (x - xbar)^2 and (y - ybar)^2 values to compute the standard deviations of height and weight respectively. Correlation, as we know, has been defined as covariance of X and Y divided by standard deviations of X and Y. This can also be represented as summation over (x - xbar) multiplied to (y - ybar) divided by the square root of summation over sum of squared deviations, that is (x - xbar)^2, multiplied to the square root of summation over (y - ybar)^2, that is the sum of squared deviations for y. Now let's find out values to put into this formula. First, we find out the overall sum of height to get the mean of height, which is 5.14. Similarly, we get the sum of weight to get the mean of weight as 50. We now get the summation over (x - xbar) multiplied to (y - ybar) to get the numerator for the formula. Then we compute (x - xbar)^2 summation and (y - ybar)^2, that is the sum of squared deviation of x and y respectively. Now, we put in the values in this final correlation formula to get a correlation value of 0.889. This indicates that height and weight have a positive relationship. It is evident that as height grows, weight also increases.

In this module, we will be talking about expectation and variance. So the expected value, or we can say the mean of a given variable that we can denote by X, is a discrete random variable where it is a weighted average of the possible values that X can take, and each value is going to be according to the probability of that specific event occurring. So usually the expected value of X is denoted by a simple formula where we can define the expectation based on the X parameter, which is going to be the sum of each possible outcome multiplied by the probability of the outcome occurring. So, in more concrete terms, the expectation is what we would expect the outcome of an experiment to be on average. We can take an example for the coin. If a coin is being tossed 10 times, then one is most likely to get five heads and five tails. The same logic can be discussed if we talk about another example of rolling a die. So there are six possible outcomes when you roll a die: 1, 2, 3, 4, 5, 6, and each of these has a probability of 1/6 of occurring. So we can say that the expectation is going to be 1 multiplied by the probability of that happening, which is going to be (1 x 1/6) + (2 x 1/6) + (3 x 1/6) + (4 x 1/6) + (5 x 1/6) + (6 x 1/6), and that is going to give us 3.5 as an output. The expected value is 3.5. So if you think about it, 3.5 is halfway between the possible values that I can take, and this is what we should have expected.

Next, we talk about the concept of variance. So the variance of a random variable allows us to know something about the spread of the possible values of the variable. So for a discrete random variable X, the variance of X is going to be denoted by using a simple formula that is going to be var = E(X - M)^2, where M is basically the expected value of the expectation of X. So this is more like a standard deviation of X, which can also be represented by using this formula. So the variance does not behave in the same way as expectation when we multiply and add constants to random variables. So now there are two different types of variance that we can have a fair understanding of. First of all, we have low variance, and then we have high variance. So low variance simply means that there is a small variation in the prediction of the target function with changes in the training data set, and at the same time, high variance, as we can see here, high variance shows a large variation in the prediction of the target function with changes in the training data set. So a model that shows high variance learns a lot and performs well with the training data set, and it does not generalize well with the unseen data set, and that's why, as a result, such a model gives good results with the training data set but shows high error rates on the test data set, and since the high variance, a model learns too much from the data set, it leads to an overfitting of the model. So a model with high variance will be having a couple of issues like it may lead to overfitting, or it may also lead to an increase in model complexities.

Next, we have skewness. So skewness, in simple terms, is basically a measure of asymmetry of a distribution. So a distribution is asymmetrical when its left and right sides are not the mirror images. Right now, this is a mirrored image, and a distribution can have right (positive) or we can say negative, or it can have zero skewness. So right-skewed, in this scenario, is basically the distribution is longer on the right side of its peak, and a left-skew distribution is going to be, we can say, where it is longer on the left side. So we can see we have this one as a part of the right side; it is more elongated towards the right side, and this one is more elongated towards the left side. So we can think of skewness in terms of tails. A tail is the long, tapering end of a distribution. So it simply indicates that there are observations at one end of the distribution, but that they are relatively infrequent. So a right-skew distribution has a long tail on the right side, as you can see here. So the number of supports observed. Let's say we have data on a per-year basis. So again, we can have more skewness towards the right side where data is being dropped as we continue to increase the number of years. For example, we may have high sales towards the beginning of the year, suppose in 2022, but again, as we proceed to 2023, second half, we are seeing the dip in performance. So that is rightly skewed, and the same way, let's suppose if we started with the sales figure, it was really less in, suppose 2002, but again, as we proceeded to 2023, now our sales have been gradually increasing, so it's more like skewed towards the left section as a part of negative skew.

Next, we have kurtosis. So kurtosis is basically a measure of the tailness of a distribution. So tailness is how often the outliers occur, and acts as kurtosis is the tailness of the distribution related to a normal distribution. So a distribution with medium kurtosis is called as mesokurtic. A distribution with low kurtosis, like this one, this is called as the platykurtic, and then a distribution with high kurtosis, like this one, this is called as the leptokurtic. So tails here, they are tapering ends on either side of a distribution like this. So they represent the probability or the frequency of values that are extremely high or extremely low to the mean. In other words, tails here represent how often the outliers occur. So there are three types of kurtosis. We have platykurtic, which is negative, leptokurtic, which is positive towards the upper end, and then we have mesokurtic, which is a normal distribution. So mesokurtic is the medium tail. So normal distributions, they have a kurtosis of three. So any distribution with a kurtosis of an approx value of three is going to be mesokurtic. And kurtosis is described in terms of excess kurtosis, which is kurtosis - 3. And since normal distributions they have a kurtosis of three, excess kurtosis makes comparing a distribution's kurtosis to a normal distribution even easier.

Introduction to probability. Probability theory. Probability is a measure of the likelihood that an event will occur. Let's consider an example of a coin toss where the chances of getting heads on a coin are 1/2 or 50%. The probability of each given event is between zero and one, both inclusive. The sum of an event's cumulative probability cannot be greater than one. Hence, the probability of an event x lies between zero and one. This means that the integral of the probability of distribution over X equals to 1.

Conditional probability. Conditional probability of any event A is defined as the probability of occurrence of A given that event B has previously occurred. Conditional probability of event A given B can be estimated as probability of (A intersection B), that is the probability of both A and B happening together, divided by the probability of B. It is also written as that probability of (A intersection B) equals to probability of A given B multiplied to probability of B.

Let's consider an example. In a coin toss, we are doing a two-coin flip. Coin one gets heads, tails, heads, and tails in subsequent flips, while coin two gets tails, heads, heads, and tails in the subsequent flips. Now, the probability that coin one will get a head is 2 out of 4. While the probability that coin two will get heads is again 2 out of 4. The probability that both coin one and coin two will have a heads is just one out of the four flips. Hence, the probability that coin one will get heads given that coin 2 is already heads can be computed as probability of (coin one heads intersection coin two heads), that is 1/4, divided by probability of (coin two heads), that's a given, that is 2/4, which is going to be 0.5 or 50%.

Bayes' theorem. Bayes' theorem calculates the conditional probability of an event based on its prior probabilities. Basically, Bayes' theorem incorporates the prior probability distribution to predict the posterior probabilities. Bayes' theorem for conditional probability can be expressed as probability of A given B equals probability of B given A divided by probability of B multiplied to probability of A. Bayes' theorem allows updating the probability values by using new information or evidence. Here, probability of A is known as prior probability; that is, the probability of an event before any new data is collected. Probability of A given B is known as the posterior probability. It is the revised probability of an event occurring after taking into consideration the new information. Probability of B given A is known as the likelihood, and probability of B is the probability of observing an evidence B model.

An example. Consider an example for calculating the likelihood of having diabetes based on the frequency of fast food consumption. Here is the observed data. Let's say the fast food audience is 20%. Diabetes prevalence is 10%, and 5% is fast food and diabetes. The chances of diabetes given fast food, that is the conditional probability of D given B, can be calculated as probability of (diabetes and fast food together) divided by probability of fast food. That means 5% divided by 20%, that equals 25%. A defined analysis can state eating fast food increases the chance of having diabetes by 25%.

The multiplication rule of probability. If events A and B are statistically independent, then probability of (A intersection B) can be given as probability of A given B multiplied to probability of B. However, probability of (A intersection B) is also given as probability of A multiplied to probability of B. Here, probability of A given B equals to probability of A when we assume that probability of B is non-zero. Similarly, probability of B equals probability of B given A, assuming probability of A is non-zero.

Chain rule of probability. Joint probability distributions over many random variables can be reduced into conditional distributions over a single variable. It can be expressed as probability of (X1, X2, ..., Xn) equals probability of X1 intersection probability of Xi given probability of (X1, ..., Xi-1). For example, the joint probability of A, B, and C can be given as probability of A given (B, C) multiplied to probability of B given C multiplied to probability of C.

Logistic sigmoid. The logistic function is a type of sigmoid function that aims to predict the class to which a particular sample belongs. Its outcome is a discrete binary value, a probability between zero and one. The logistic sigmoid is a useful function that follows the S-curve. It saturates when the input is very large or very small. Logistic sigmoid is expressed as sigma(x) = 1 / (1 + e^-x). The logistic sigmoid can be expressed as sigmoid function of x is given as 1 / (1 + e^-x), where e is Euler's number.

Gaussian distribution. The Gaussian distribution is a type of distribution in which data tends to cluster around a central value with little or no bias to the left or right. It is often referred to as a normal distribution. In the absence of prior information, the normal distribution is frequently a fair assumption in machine learning equations. The formula for calculating Gaussian distribution is described as the normal distribution of X. That is, the function of x given mean as mu and variance is sigma^2 can be calculated as 1 / (sigma * sqrt(2*pi)) * e^(-(x - mu)^2 / (2*sigma^2)), where mu is the mean or peak value, which also is the expected value of x. Sigma is the standard deviation. Sigma^2 is the variance. A standard normal distribution has a mean of zero and a standard deviation of one. Gaussian distribution can be univariate, which describes the distribution of a single variable X. It can also be multivariate, where it can just be used to describe the distribution of several variables. It is represented in 3D or nD formats.

Law of large numbers. Now let's talk about the law of large numbers. The law of large numbers states that an observed sample average from a large sample will be close to the true population average, and that it will get closer in a larger sample. So the law of large numbers does not guarantee that a given sample, especially a small sample, will reflect the true population characteristics, or that a sample that does not reflect the true population will be balanced by a subsequent sample. This is for the law of large numbers to express the relationship between scale and growth rate. So there are multiple examples through which we can understand, and it is widely used in statistical analysis in working with the central limit theorem in terms of business growth. So there are multiple real-time setups in which these are going to be used. So if you talk about tossing a coin, so tossing a coin a number of times will give us...

Two different types of outcomes are possible. The result will spread evenly between heads and tails, and the expected average value is going to be half. That means 50 times tails and 30 heads. But again, if you toss a coin 1,000 times, then the result can be different because out of 1,000, let's say 850 times it has been heads and only 150 times it has been tails, and so on. So that's why the possibility of one event occurring is going to be changed in large sample sets as compared to small sample sets, as in, let's say, 10 times. So the number of heads and tails is unbalanced for a lower number of trials. So we can see it is unbalanced. But again, as soon as we toss more coins, the result leans towards the balance value, or we can see the observed averages.

Next, we have the p-value. So a p-value is basically a number calculated from a statistical test that describes how likely we are to have found a particular set of observations if the null hypothesis were true. So p-values are used in hypothesis testing to help decide whether to reject the null hypothesis. And the smaller the p-value, the more likely we are to reject the null hypothesis. So we have a term called the null hypothesis. All statistical tests have a null hypothesis. For most tests, the null hypothesis is that there is no relationship between our variables or that there is no difference among groups. For example, in a two-tailed t-test, the null hypothesis is that the difference between two groups is going to be zero. So the p-value is going to tell us how likely it is that our data could have occurred under the null hypothesis. It is done by calculating the likelihood of a test statistic, which is the number calculated by a statistical test using our data. So the p-value tells us how often we would expect to see a test statistic as extreme or more extreme than the one calculated by a statistical test if the null hypothesis of the test was true.

So there are multiple limitations as well. First, the results can be significant, but they may not be practical, as we have compared it. It can be based on multiple hypotheses for a game or a healthcare test. If the test is going to be positive or not, it may show even values of the effect of a variable but not the magnitude in real life. What exactly is going to be the application of a drug test being failed in a pharma company? Therefore, it is recommended to use confidence levels in addition to p-values to quantify, or we can say to give a solid figure to, the reserve which we are going to get. The p-values are interpreted as supporting or refuting the alternative hypothesis. So a p-value can only tell you whether or not the null hypothesis is supported. It cannot tell us whether our alternative hypothesis is true or why. So the risk of rejecting the null hypothesis is often higher than the p-value, especially when we are looking at a single study or when using small sample sizes. This is because the smaller the frame of reference, the greater the chance that we stumble across a statistically significant pattern completely by accident.

Key takeaways. Key takeaways. Probability and statistics structure the premise of the data. The data helps in anticipating the future or gauging in view of the past patterns of information. The central tendency is a single value that helps describe the data by identifying its central positions. The mean, median, and mode are the measures of central tendencies. The distribution where the data tends to be around a central value with a lack of bias or minimal bias towards the left or right is called a Gaussian distribution.

LLMs. If you ever wondered how machine learning can now understand and generate humanlike text, you are in the right place. From chatbots like ChatGPT to AI assistants that power search engines, LLMs are transforming how we interact with technology. One of the most exciting advancements in this space is Google's Gemini or OpenAI's ChatGPT, large language models designed to push the boundaries of what AI can achieve. In this video, we will explore what LLMs are, how they work, and why models like Gemini are critical for the future of AI. Google Gemini is part of a new wave of AI models that are smarter, faster, and more efficient. It is designed to understand context better, offer more accurate responses, and integrate deeply into services like Google Search and Google Assistant, providing more humanlike interactions. So we will break down the science behind LLMs, including their massive training data sets, transformer architecture, and how models like Gemini use deep learning innovations to change industries. Plus, we will compare Google Gemini to other popular LLMs such as OpenAI's ChatGPT models, showing how each of these technologies is used to power chatbots, virtual assistants, and other AI-driven applications. By the end of this video, you will have a clear understanding of how large language models like Gemini work, their key features, and what they mean for the future of AI. Don't forget to like, subscribe, and hit the bell icon to never miss any update from Simply Learn. So, without any further ado, let's get started.

So, what are large language models? Large language models like ChatGPT-4, Generative Pre-trained Transformer 4, GPT, and Google Gemini are sophisticated AI systems designed to comprehend and generate humanlike text. These models are built using deep learning techniques and are trained on vast data sets collected from the internet. They leverage self-attention mechanisms to analyze relationships between words or tokens, allowing them to capture context and produce coherent, relevant responses. LLMs have significant applications, including powering virtual assistants, chatbots, content creation, language translation, and supporting research and decision-making. Their ability to generate fluent and contextually appropriate text has advanced natural language processing and improved human-computer interaction.

So now let's see what large language models are used for. Large language models are utilized in scenarios with limited or no domain-specific data available for training. These scenarios include both few-shot and zero-shot training approaches, which rely on the model's strong inductive bias and its capability to derive meaningful representations from a small amount of data or even no data at all.

So now let's see how large language models are trained. Large language models typically undergo pre-training on a broad, all-encompassing data set that shares statistical similarities with the data set specific to the target task. The objective of pre-training is to enable the model to acquire high-level features that can later be applied during the fine-tuning phase for specific tasks. So there are some training processes of LLMs which involve several steps. The first one is text pre-processing. The textual data is transformed into a numerical representation that the LLM model can effectively process. This conversion may involve techniques like tokenization, encoding, and creating input sequences. The second one is random parameter initialization. The model's parameters are initialized randomly before the training process begins. The third one is input numerical data. The numerical representation of the text data is fed into the model for processing. The model's architecture, typically based on transformers, allows it to capture the conceptual relationships between the words or tokens. The fourth one is loss function calculation. A loss function calculation measures the discrepancy between the model's prediction and the actual next word or token in a sequence. The LLM model aims to minimize this loss during training. The fifth one is parameter optimization. The model's parameters are adjusted through optimization techniques. This involves calculating gradients and updating the parameters accordingly, gradually improving the model's performance. The last one is iterative training. The training process is repeated over multiple iterations or epochs until the model's output achieves a satisfactory level of accuracy on the given task or data set. By following this training process, large language models learn to capture linguistic patterns, understand context, and generate coherent responses, enabling them to excel at various language-related tasks.

The next topic is how large language models work. So large language models leverage deep neural networks to generate output based on patterns learned from the training data. Typically, a large language model adopts a transformer architecture, which enables the model to identify relationships between words in a sentence irrespective of their position in the sequence. In contrast to RNNs that rely on recurrence to capture token relationships, transformer neural networks employ self-attention as their primary mechanism. Self-attention calculates attention scores that determine the importance of each token with respect to the other tokens in the text sequence, facilitating the modeling of intricate relationships within the data.

Next, let's see applications of large language models. Large language models have a wide range of applications across various domains. Here are some notable applications: Natural language processing (NLP). Large language models are used to improve natural language understanding tasks such as sentiment analysis, named entity recognition, text classification, and language modeling. Chatbots and virtual assistants. Large language models power conversational agents, chatbots, and virtual assistants, providing more interactive and humanlike user interaction. Machine translation. Large language models have been used for automatic language translation, enabling text translation between different languages with improved accuracy. Sentiment analysis. LLMs can analyze and classify the sentiment or emotion expressed in a piece of text, which is valuable for market research, brand monitoring, and social media analysis. Content recommendation. These models can be employed to provide personalized content recommendations, enhancing user experience and engagement on platforms such as news websites or streaming services. These applications highlight the potential impact of large language models in various domains for improving language understanding and automation.

There are a lot of areas where data science can be used. One of the very common ones is fraud detection or fraud prevention. There are a lot of fraudulent activities or transactions, primarily on the internet. It's very easy to commit fraud, and therefore we can use data science to either prevent or detect fraud. There are certain algorithms, machine learning algorithms, that can be used, like, for example, some outlier techniques, clustering techniques, that can be used to detect and prevent fraud as well.

So who is a data scientist, rather? It is actually a very generic role that defines somebody who is working with data as a data scientist. But there can be very specific activities, and the roles can be much more specific. What exactly a person does within the area of data science can be much more specific. But broadly, anybody working in the area of data science is known as a data scientist.

So what does a data scientist do? These are some of the activities: data acquisition, data preparation, data mining, data modeling, and model maintenance. We will talk about each of these in great detail, but at a very high level, the first step obviously is to get the raw data, which is known as data acquisition. It can be all kinds of formats and could be from multiple sources, but obviously that raw data cannot be used as it is for performing data mining activities or data modeling activities. So the data has to be planned and prepared for use in the data models or in the data mining activity. So that is data preparation. Then we actually do the data mining, which can also include some exploratory activities. And then if we have to do stuff like machine learning, then you need to build a machine learning model and test the model, get insights out of it, and then if the model is fine, you deploy it, and then you need to maintain the model because over a period of time it is possible that you need to tweak the model because of changes in the process or changes in the data and so on. So that all comes under model maintenance.

So let's take a deeper look at each of these activities. Let's start with data acquisition. So the stage of data acquisition is basically where the data scientist will collect raw data from all possible sources. This could be typically an RDBMS, which is a relational database, or it can also be a non-RDBMS, or could be flat files or unstructured data, and so on. So we need to bring all that data from different sources if required. We need to do some kind of homogeneous formatting so that it all fits into, at least from a format perspective, it looks homogeneous. So that may be requiring some kind of transformation. Very often this is loaded into what is known as a data warehouse. So this can also be sometimes referred to as ETL or extract, transform, and load. So a data warehouse is like a common place where the data from different sources is brought together so that people can perform data science activities like reporting or data mining or statistical analysis and so on. So data from various sources is put in a centralized place, which is known as a data warehouse. So that is also known as ETL. And in order to do this, data scientists can take help of some ETL tools. There are some existing tools that a data scientist can take help of, like, for example, DataStage, Talend, or Informatica. These are pretty good tools for performing these ETL activities and getting the data.

The next stage, now that you have the raw data into a data warehouse, you still probably are not in a position to straight away use this data for performing the data mining activities. So that is where data preparation comes into play, and there are multiple reasons for that. One of them could be that the data is dirty. There are some missing values and so on and so forth. So a lot of time is actually spent in this particular stage. So a data scientist spends a lot of time, almost 60 to 70% of the time, in this part of the project or the process, which is data preparation. So there are again within this, there can be multiple sub-activities, starting from, let's say, data cleaning. You will probably have missing values; the data has some columns where the values are missing or the values are incorrect; there are null values and so on and so forth. So that is basically the data cleaning part of it. Then you need to perform certain transformations, like, for example, normalizing the data and so on, or you could probably have to modify categorical values into numerical values and so on and so forth. So these are transformational activities. Then we may have to handle outliers. So the data could be such that there are a few values which are way beyond the normal behavior of the data, for whatever reason, either people have keyed in wrong values or for some reason some of the values are completely out of range. So those are known as outliers. So there are certain ways of handling these outliers and detecting and handling these outliers. So this is a part of what is known as exploratory analysis. So you quickly explore the data to find out are there. So, and you can use visual tools like plots and identify what are the outliers and see how we can get rid of the outliers and so on. Then the next part could be data integrity. Data integrity is to validate, for example, if there are some primary keys that all the primary keys are populated; if there are some foreign keys, then at least most of the foreign keys should be populated, and otherwise, when we are trying to query the data, you may get wrong values and so on. So that is the data integrity part of it, and then we have what is known as data reduction. Sometimes we may have duplicate values; we may have columns that may be duplicated because they're coming from different sources; the same values are there and so on. So a lot of this can be done using what is known as data reduction, and thereby you can reduce the size of the data drastically because very often this could be redundant data which can be removed and so on.

So let's take a look at what are the various techniques that are used for data cleaning. So we need to ensure that the data is valid and it is consistent and uniform and accurate. So these are the various parameters that we need to ensure as part of the data cleaning process. Now what are the techniques that are used for data cleaning? So we will see what each of these are in this particular case, and so what is the data set that we have? We have data about a bank and its customer details. So let's take an example and see how we go about cleaning the data. And in this particular example, we're assuming we are using Python. So let's assume we loaded this data, which is the raw file CSV. This is how the customer data looks like, and we will see, for example, we take a closer look at the geography column; we will see that there are quite a few blank spaces. So how do we go about when we have some blank spaces? Or if it is a string value, then we put an empty string here or we just use a space or empty string. If they are numerical values, then we need to come up with a strategy. For example, we put the mean value. So wherever it is missing, we find the mean for that particular column. So in this case, let's assume we have credit score, and we see that quite a few of these values are missing. So what do we do here? We find the mean for this column for all the existing values, and we found that the mean is equal to 638.66. So we kind of write a piece of code to replace wherever there are blank values. NaN is basically like null, and we just go ahead and say fill it with the mean value. So this is the piece of code we are writing to fill it. So all the blanks or all the null values get replaced with the mean value. Now one of the reasons for doing this is that very often if you have some such situation, many of your statistical functions may not even work. So that's the reason you need to fill up these values or either get rid of these records or fill up these values with something meaningful. So this is one mechanism, which is basically using a mean. There are a few others as we move forward. We can see what are the other ways. For example, we can also say that any missing value in a particular row, if even one column the value is missing, you just drop that particular row or delete all rows where even a single column has missing values. So that is one way of dealing. Now the problem here can be that if a lot of data has, let's say, one or two columns missing, and we drop many such rows, then overall you may lose out on, let's say, 60% of the data has some value or the other missing; 60% of the rows, then it may not be a good idea to delete all the rows like in that manner because then you're losing pretty much 60% of your data, therefore your analysis won't be accurate. But if it is only 5 or 10%, then this will work. Another way is only to drop values where, or rather drop rows where all the columns are empty, which makes sense because that means that record is of really no use because it has no information in it. So there can be some situations like that. So we can provide a condition saying that drop the records where all the columns are blank or not applicable. We can also specify some kind of a threshold. Let's say you have 10 or 20 columns in a row. You can specify that maybe five columns are blank or null, then you drop that record. So again, we need to take care that in such a situation, the amount of data that has been removed or excluded is not large. If it is like maybe 5%, maximum 10%, then it's okay. But by doing this, if you're losing out on a large chunk of data, then it may not be a good idea. You need to come up with something better. What else we need to do next is so the data preparation part is done. So now we get into the data mining part. So what exactly we do in data mining? Primarily, we come up with ways to take meaningful decisions. So data mining will give us insights into the data what is existing there, and then we can do additional stuff like maybe machine learning and so on to get perform advanced analytics and so on. So one of the first steps we do is what is known as data discovery, which is basically like exploratory analysis. So we can use tools like Tableau for doing some of this. So let's just take a quick look at how we go about that. So Tableau is an excellent data mining or actually more of a reporting or a BI tool, and you can download a trial version of Tableau at tableau.com, or there is also Tableau Public, which is free, and you can actually use and play around. However, if you want to use it for enterprise purposes, then it is commercial software. So you need to purchase a license and you can then run some of the data mining activities. Let's say your data source, your data is in some Excel sheet. So you

Can select the source as Microsoft Excel or any other format, and the data will be brought into the Tableau environment. Then it will show you what is known as dimensions and measures. Dimensions are all the descriptive columns. Tableau is intelligent enough to actually identify these dimensions and measures. Measures are the numerical values. As you can see here, customer ID, gender, geography—these are all dimensions, non-numerical values—whereas age, balance, credit score, and so on are numeric values. They come under measures.

So you've got your data into Tableau, and then you want to, let's say, build a small model and solve a particular problem. What is the problem statement? Let's say we want to analyze why customers are leaving the bank, which is known as exit, and we want to analyze and see what are some of the factors for exiting the bank. We want to, let's assume, consider these three—gender, credit card, and geography—as criteria and analyze if these are in any way impacting or have some bearing on the customer exiting or the customer exit behavior. Okay.

Let's use Tableau, and very quickly we will be able to find out how these parameters are affecting. This is our customer data. From our Excel sheet, we have a data set of about 10,000 rows, and we want to find out what the criteria are. Let's start with gender. Let's say we want to first use gender as a criteria. Tableau really offers an easy drag-and-drop kind of mechanism, making it really easy to perform this kind of analysis. Exited says whether the customer has exited or not. It has a value of zero and one, and then, of course, you have gender and so on. We will take these two and simply drag and drop.

Okay. So exited, and then we will put gender. If we drag and drop into the analysis side of Tableau, we are showing male and female as two different columns here, zero for people who did not exit and one for people who exited, and that is color-coded. The blue color means people who did not exit, and this yellow color means people who did exit.

Now, if we pull the data here to create bar graphs, this is how it would look. Yellow is who exited, and for the male, only 16.45% have exited. We can also draw a reference line that will help us or even provide aliases. These are a lot of fancy things provided by Tableau. You can create aliases so that it looks good rather than basic labels, and you can also add a reference line. From here, we can make out that, on average, female customers exit more than male customers. That is what we are seeing here, on average. We have analyzed based on gender. We do see that there is some difference in the male and female behavior.

Now let's take the next criteria, which is the credit card. Let's see if having a credit card has any impact on the customer exit behavior. Just like before, we drag and drop the credit card—the "has credit card" column—and then we will see that there is pretty much no difference between people having a credit card and not having a credit card. 20.81% of people who have no credit card have exited, and similarly, 20.18% of people who have a credit card have also exited. The credit card is not having much of an impact. That's what this piece of analysis shows.

Last, we will check how geography is impacting. Once again, we can drag and drop the geography column onto this side. If we see here, there are geographies like—I think there are about three geographies—like France, Germany, and Spain, and we see that there is some kind of impact with the geography as well. What we derive from this is that the credit card is—we can ignore the credit card variable or feature from our analysis because that doesn't have any impact—but gender and geography, we can keep and do further analysis.

What are some of the advantages of data mining? Bit more detailed analysis can help us in predicting future trends, and it also helps in identifying customer behavior patterns. You can make informed decisions because the data is telling you or providing you with some insights, and then you make a decision based on that. If there is any fraudulent activity, data mining will help in quickly identifying such fraud as well, and, of course, it will also help us in identifying the right algorithm for performing more advanced data mining activities like machine learning and so on.

The next activity—now that we have the data, we have prepared the data, and performed some data mining activity—the next step is model building. Let's take a look at model building. What is model building? If we want to perform a more detailed data mining activity, like maybe perform some machine learning, then you need to build a model. How do you build a model? First, you need to select which algorithm you want to use to solve the problem at hand, and also what kind of data is available and so on and so forth. You need to make a choice of the algorithm, and based on that, you go ahead and create a model, train the model, and so on.

Machine learning is, at a very high level, classified into supervised and unsupervised. If we want to predict a continuous value—it could be a price, a temperature, a height, a length, or things like that—those are continuous values, and if you want to find some of those, then you use techniques like regression: linear regression, simple linear regression, multiple linear regression, and so on. These are the algorithms.

On the other hand, there will be situations where you need to perform unsupervised learning. In unsupervised learning, you don't have any historical labeled data to learn from. That is when you use unsupervised learning. Some of the algorithms in unsupervised learning are clustering. K-means clustering is the most common algorithm used in unsupervised learning.

Similarly, in supervised learning, if you want to perform some activity on categorical values—like, for example, it is not measured but it is counted—like you want to classify whether this image is a cat or a dog, whether you want to classify whether this customer will buy the product or not, or you want to classify whether this email is spam or not spam. These are examples of categorical values, and these are examples of classification. Then you have algorithms like logistic regression, K-nearest neighbor or KNN, and support vector machine. These are some of the algorithms that are used in this case. Similarly, in unsupervised learning, if you need to perform on categorical values, you have some algorithms like association analysis and hidden Markov model.

To understand this better, let's take an example and take you through the whole process, and then we will also see how the code can be written to perform this. Now let's take our example here where we want to perform supervised learning, which is basically we want to do a multiple linear regression, which means there are multiple independent variables, and then you want to perform a linear regression to predict a certain value.

In this particular example, we have world happiness data. This is data about the happiness quotient of people from various countries, and we are trying to predict and see how our model will perform. What is the question that we need to ask? First of all, how to describe the data, and then, can we make a predictive model to calculate the happiness score? Based on this, we can then decide on what algorithm to use and what model to use and so on.

Variables that are available or used in this model: This is a list of variables that are available. There is a happiness rank, happiness score (the happiness score is more like an absolute value, whereas rank is what the ranking is), which country we are talking about, within that country which region, what kind of economy, whether the family—which family—and health details, freedom, trust, generosity, and so on and so forth. There are multiple variables that are available to us, and the specific details probably are not required, and there can be—in another example—the variables can be completely different. We don't have to go into the details of what exactly these variables are, but it's just enough to understand that we have a bunch of these variables, and now we need to use either all or some of these variables (which we also sometimes refer to as features), and then we need to build our model and train our model.

Let's assume we will use Python to perform this analysis or perform this machine learning activity. I will actually show you in our lab, in a little bit, this whole thing. We will run the live code. But quickly, I will run you through the slides, and then we will go into the lab.

What are we doing here? First, we need to import a bunch of libraries in Python which are required to perform our analysis. Most of these are for manipulating the data, preparing the data, and then scikit-learn or sklearn is the library which you will use actually for this particular machine learning activity, which is linear regression. So we have NumPy, we have Pandas, and so on and so forth. All these libraries are imported, and then we load our data, and the data is in the form of a CSV file, and there are different files for each year. So we have data for 2015, 16, and 17. So we will load this data and then combine them, concatenate them, to prepare a single data frame. Here we are making an assumption that you are familiar with Python. So it becomes easier if you are familiar with the Python programming language, or at least some programming language, so that you can at least understand by looking at the code.

So we are reading the file, each of these files for each year, and this is basically—we are creating a list of all the names of the columns we will be using later on; you will see in the code. So we have loaded 2015, then 2016, and then also 2017. So we have created three data frames, and then we concatenate all these three data frames. This is what we are doing here. Then we identify which of these columns are required. Which, for example, some of the categorical values—do we really need them? We probably don't. Then we drop those columns so that we don't unnecessarily use all the columns and make the computation complicated. We can then create some plots using the Plotly library, and it has some powerful features, including creation of maps and so on, just to understand the pattern—the happiness quotient—or how the happiness is across all the countries. It's a nice visualization; we can see each of these countries, how they are in terms of their happiness score. This is the legend here. The lighter colored countries have lower ranking, and these are the lower ranking ones, and these are higher ranking, which means that the ones with these dark colors are the happiest ones. As you can see here, Australia and maybe beside the US and so on are the happiest ones.

The other thing that we need to do is the correlation between the happiness score and happiness rank. We can find a correlation using a scatter plot, and we find that, yes, they are kind of inversely proportional, which is obvious. If the score is high, the happiness score is high, then they are ranked number one. For example, the highest is scored as number one. That's the idea behind this. The happiness score given here and the happiness rank is actually given here. They are inversely proportional because the higher the score, the absolute value of the rank will be lower. Number one has the highest value of the score and so on. So they are inversely correlated, but there is a strong—what this graph shows is that there is a strong correlation between happiness rank and happiness score. And then we do some more plots to visualize this. We determined that probably rank and score are pretty much conveying the same message, so we don't need both of them. So we will drop one of them. That is what we are doing here. So we drop the happiness rank, and similarly... This is one example of how we can remove some columns which are not adding value. We will see in the code as well how that works.

Moving on. This is a correlation between pretty much each of the columns with the other columns. This is a correlation you can plot using the plot function, and we will see here that, for example, happiness score and happiness score are correlated—strongest correlation—right, because every variable will be highly correlated to itself. So that's the reason—so the darker the color is, the higher the correlation, and so the—and correlation in numerical terms goes from 0 to 1. One is the highest value, and it can only be between 0 and 1. Correlation between two variables can only have a value between 0 and 1. The numerical value can go from 0 to 1, and 1 here is dark color, and 0 is kind of dark, but it is blue color. From red, it goes down; the dark blue color indicates pretty much no correlation. So the—from this heat map, we see that happiness and economy and family are probably also—health probably—are the most correlated, and then it keeps decreasing after freedom, kind of keeps decreasing and coming to pretty much zero.

That is a correlation graph, and then we can probably use this to find out which are the columns that need to be dropped, which do not have very high correlation, and we take only those columns that we will need. This is the code for dropping some of the columns. Once we have prepared the data, when we have the required columns, then we use scikit-learn to actually split the data. First of all, this is a normal machine learning process. You need to split the data into training and test data sets. In this case, we are splitting into 80/20. So 80% is the training data set, and 20% is the test data set. That's what we are doing here. So we use the train_test_split method or function. So you have all your training data in X_train, the labels in Y_train. Similarly, X_test has the test data, the inputs, whereas the labels are in Y_test. That's how—and this value, whether it is 80/20 or 50/50, that is all individual preference. In our case, we are using 80/20.

Then the next is to create a linear regression instance. This is what we are doing. We're creating an instance of linear regression, and then we train the model using the fit function. We are passing x and y, which is the x value and the label data—regular input and the label data—label information. Then we do the test; we run or we perform the evaluation on the test data set. This is what we are doing with the test data set, and then we will evaluate how accurate the model is, and using the scikit-learn functionality itself. We can also see what are the various parameters and what are the various coefficients because in linear regression you will get like an equation of like a straight line—y = β₀ + β₁x₁ + β₂x₂—those β₁, β₂, β₃ are known as the coefficients, and β₀ is the intercept. After the training, you can actually get this information of the model—what is the intercept value, what are the coefficients, and so on—by using these functions. Let's take a quick look into the lab and take a look at our code.

Okay. This is my lab. This is my Jupyter Notebook where I have the actual code, and I will take you through this code to run this linear regression on the world happiness data. So we will import a bunch of libraries—NumPy, Pandas, Plotly, and so on, also... yeah, scikit-learn, that's also very important. So that's the first step. Then I will import my data, and the data is in three parts. There are three files, one for each year: 2015, 2016, and 2017. And it is a CSV file. So I've imported my data. Let's take a quick look at the data. This is how it looks. We have the country, region, happiness rank, and then happiness score. There are some standard errors, and then what is the per capita family and so on. And then we will keep going. We will create a list of all these column names we will be using later. So for now, I will run this code. No need for major explanation at this point. We know that some of these columns probably are not required. So you can use this drop functionality to remove some of the columns which we don't need—like, for example, region and standard error will not be contributing to our model. So we will drop those values out here. So we use the drop, and then we created a vector with these names—column names. That's what we are passing here. Instead of giving the names of the columns here, we can pass a vector. So that's what we are doing. So this will drop from our data frame; it will remove region and standard error—these two columns. Then the next step, we will read the data for 2016 and also 2017, and then we will concatenate this data. So let's do that. So we have now a data frame called happiness, which is a concatenation of all three files. Let's take a quick look at the data now. Most of the unwanted columns have been removed, and you have all the data in one place for all three years. And this is how the data looks. And if you want to take a look at the summary of the columns, you can say describe, and you will get this information. For example, for each of the columns, what is the count? What's the mean value, standard deviation—especially the numeric values, okay?—not the categorical values. So this is a quick way to see how the data is, and initial little bit of exploratory analysis can be done here. So what is the maximum value? What's the minimum value and so on for each of the columns.

All right. So then we go ahead and create some visualizations using Plotly. So let us go and build a plot. So if we see here now, this is the relation—correlation—between happiness rank and happiness score. This is what we have seen in the slides as well. We can see that there is a tight correlation between them. Only thing is it is inverse correlation, but otherwise they are very tightly correlated, which also says that they both probably provide the same information. So there is not much value added. So we'll go ahead and drop the happiness rank as well from our columns. So that's what we're doing here. And now we can do the creation of the correlation heat map. Let us plot the correlation heat map to see how each of these columns is correlated to the others, and, as we have seen in the slides, this is how it looks. So happiness score is very highly correlated. This is the legend we have seen in the slide as well. So blue color indicates pretty much zero or very low correlation. Deep red color indicates very high correlation, and the value—correlation is a numeric value—and the value goes from 0 to 1. If the two items or two features or columns are highly correlated, then they will be as close to 1 as possible, and two columns that are not at all correlated will be as close to 0 as possible. So that's how it is. For example, here, happiness score and happiness score—every column or every feature will be highly correlated to itself. So it is like between them; there will be a correlation value will be 1. So that's why we see deep red color. But then others...

are, for example, with higher values are economy, and then health, and then maybe family and freedom. So these are generosity and trust are not very highly correlated to happiness score. So that is uh one quick exploratory analysis we can do, and uh therefore we can drop the country and the happiness rank because they also again don't have any major impact on the analysis, on our analysis.

So now we have prepared our data. There was no need to clean the data because the data was clean. But if there were some missing values and so on, as we have discussed in the slides, we would have had to perform some of the data cleaning activities as well. But in this case, the data was clean. All we needed to do was just the preparation part. So we removed some unwanted columns and we did some exploratory data analysis. Now we are ready to perform the machine learning activity.

So we use scikit-learn for doing the machine learning. Scikit-learn is a Python library that is available for performing our uh machine learning. Once again we will import some of these libraries like pandas and numpy and also scikit-learn. First step we will do is split the data in uh 20/80 format. So you have all the test data which is 20% of the data is test data and 80% is your training data. So this test size indicates how much of it is in the what is the size of the test data; remaining which is here we are saying 0.2, therefore that means training is 80%; so training data is 80%.

All right, so we have executed that split the data, and now we create an instance of the linear regression model, so lm is our linear regression model, and we pass x and y, the training data set, and call the function fit so that the model gets trained. So now once that is done, training is done, training is completed, and now what we have to do is we need to predict the values for the test data. So the next step is using—so you see here fit will basically run the training method. Predict will actually predict the values. So we are passing the input values which is the independent variables and we are asking for the values of the dependent variable which is which we are capturing in y_prime, and we use the predict method here lm.predict. So this will give us all the predicted y values, and remember we already have y_test has the actual values which are the labels so that we can use these two to compare and find out how much of it is error.

So that's what we are doing here. We are trying to find the difference between the predicted value and the actual value. Y_test is the actual value for the test data and Y_predict is the predicted value. We just found out the predicted value. So we will run that and we can do a quick check as to how the data looks. How is the difference? So in some cases it is positive, some cases it is negative, but in most of the cases I think the difference is very small. This is exponential to the power of 0.04 and so on. So looks like our model has performed reasonably well. We can now check some of the parameters of our model like the intercept and the coefficients. So that's what we are doing here. So these are the coefficients of the various parameters that we or the coefficients of the various independent variables. Okay. So these are the values. Then we can quickly go ahead and list them down as well against the corresponding independent variables. So the coefficients against the corresponding independent variable. So 1.0051 is the coefficient for economy, 0.99983 is for family, coefficient for family and health and so on and so forth. Right? So that's what this is showing.

Now we can use the readily available functionality of scikit-learn and then plot that to find some of the parameters which determine the accuracy of this model, like for example, what is the mean square error and so on. So that's what we are doing here. So let's just go ahead and run this. So you can see here that the root mean square error is pretty low, which is a good sign and uh which is one of the measures of uh how well our model is performing. We can do one more quick plot to just see how the actual values and the predicted values are looking. And once again you can see that as we have seen from the root mean square error, root mean square error is very very low. So that means that the actual values and the predicted values are pretty much matching up, almost matching up, and this plot also shows the same. So this line is going through the predicted values and the actual values, and the difference is very very low. So again this is actual data. This is one example where the accuracy is high and the predicted values are pretty much matching with the actual values. But in real life you may find that these values are slightly more scattered, and you may get the error value can be relatively on the higher side, the root mean square error.

Okay. So this was a good and quick example of uh the code to perform a data science activity or a machine learning or data mining activity. In this case we did what is known as linear regression. So let's go back to our slides and see what else is there. So we saw this. These are the coefficients of each of the features in our code. And uh we have seen the root mean square error as well. And uh with we can take a few hundred countries, certain values, and actually predict to see if how the model is performing, and I think we have done this as well. And in this case, as we have seen, pretty much the predicted values and the actual values are pretty much matching, which means our model is almost 100% accurate. As I mentioned, in real life it may not be the case, but in this particular case we have got a pretty good model, which is very good also. Subsequently we can assume that this is how the equation in linear regression, the model is nothing but an equation like y = β0 + β1x1 + β2x2 + β3x3 and so on. So this is what we are showing here. So this is our intercept, which is β0, and then we have β1 into economy value, β2 into the family value, β3 into health value and so on. So that is what is shown here.

Okay. So I think the next step, once we have the results from the data mining or machine learning activity, the next step is to communicate these results to the appropriate stakeholders. So that is what we will see here now. So how do we communicate? Usually you take these results and then either prepare a presentation or put it in a document and then show them these actionable results or actionable insights, and uh you need to find out who are your target audience and uh put all the results in context and uh maybe if there was a problem statement you need to put this results in the context of the problem statement; what was our initial goal that we wanted to achieve. So that we need to communicate here based on—you remember we started off with what is the question and what is the data and so on and then what is the answer. So we we need to put the results and then what is the methodology that we have used, all that has to be put and clearly communicated in business terms so that the people understand very well from a business perspective.

So once the model building is done, once the results are published and communicated, the last part is maintenance of this model. Now very often what can happen is the model may have to be subsequently updated or modified because of multiple reasons. Either the the data has changed, the way the data comes has changed, or the process has changed, or for whatever reason the accuracy may keep changing. Once you have a trained model, the—for example we got a very high accuracy, but then over a period of time there can be various factors which can cause that. So from time to time we need to check whether the model is performing well or not. The accuracy needs to be tested once in a while, and if required you may have to rebuild or retrain the model. So you do the assessment, you you see if it needs any tweaks or changes and then if it is required you need to probably retrain the model with the latest data that you have and then you deploy it. You build the model, train it, and then you deploy it. So that is like the maintenance cycle that you may have to take the model.

Data analyst versus data engineer versus data scientist. Which one to choose? This is one of the most popular questions asked by learners looking for a career in data and analytics. I'm sure you two would have come across these job roles in the ever-growing data science landscape. Though they all deal with data, these jobs are not the same. There are significant differences between what a data analyst, data engineer, and a data scientist does. We will look at these job roles and the differences in detail. First, let's look at some data analytics and data science trends. The analytics and data science market is thriving. Data analytics, data engineering, and data science are the key trends in today's exhilarating market. As per statista.com, the global big data analytics market revenue will grow at a CAGR of 30% with revenue reaching over 68 billion US$ by 2025. According to Technavio, the enterprise data management market is expected to increase by 64.08 billion US$ by 2025 as per marketsandmarkets.com. The big data market size is projected to grow from 162.6 billion US$ in 2021 to 273.4 billion US$ in 2026. Now another report from Research Dive says that the data science platform market is estimated to reach 224.3 billion US$ by 2026. So with so much data available and companies making huge investments to drive business insights, the job opportunities for data analysts, data engineers, and data scientists are going to increase in 2022 and over the coming years.

Now let's learn the major differences between data analyst versus data engineer versus data scientist. So who are they? A data analyst analyzes and interprets vast volumes of data in order to extract meaningful information out of it. They find solutions to a business problem and make critical business decisions. The insights provided by data analysts are important to companies that want to understand the needs of their end customers. But talking about who a data engineer is, a data engineer, on the other hand, builds infrastructure and scalable pipelines to manage the flow of data and prepare it for analysis. So basically they optimize the systems that enable data analysts and data scientists to perform their job efficiently. Data scientists are professionals who analyze and visualize existing data and use algorithms to build predictive models for making future decisions. They also engage with business leaders to understand their needs and present complex findings.

With that, let's look at the primary roles and responsibilities of these three job roles. Data analysts are responsible to collect, clean, store, and process data. They discover hidden patterns from data by performing exploratory data analysis and visualize data by creating charts and graphs. Acquiring data from primary and secondary sources is one of their key tasks. They build reports and dashboards and also maintain databases. Now talking about the roles and responsibilities of a data engineer. A data engineer performs data acquisition, the design, build, and test data as well to develop and maintain data architecture. Data engineers are tasked with testing, integrating, managing, and optimizing data from a variety of sources. So they integrate data into existing data pipelines, prepare data for modeling, and perform various ETL operations. Now talking about the roles and responsibilities of a data scientist. So data scientists develop machine learning models to identify trends in data for making decisions. They develop hypotheses and use their knowledge of statistics, data visualization, and machine learning to forecast the future for the business. Data scientists visualize data and use storytelling techniques and also write programs to automate data collection and processing.

Now move on to the skills possessed by data analysts, data engineers, and data scientists. To become a data analyst, you need to have good hands-on experience with writing SQL queries. You should have excellent Microsoft Excel skills for analyzing data. Data analysts are also good at programming, and they need to know how to visualize data, solve business problems, and possess domain knowledge. Data engineers should have a solid understanding of SQL, MongoDB, and programming. They need to have a good command of data architecture, scripting, data warehousing, and ETL. Data engineers are also good at Hadoop-based analytics. Now talking about the skills for a data scientist. So a data scientist should have experience with programming in Python and R. They should have a very good understanding of mathematics and statistics as well. Data scientists need to possess analytical thinking and data visualization skills as well. Machine learning, deep learning, and decision-making are other critical skills every data scientist should have.

Now we look at the salaries of a data scientist, a data analyst as well as a data engineer. So a data analyst in the United States earns over $70,000 per annum, while in India a data analyst can earn nearly 7,25,000 rupees per annum. A data engineer in the United States can earn over $112,500 per year, and in India you can earn over 9 lakh rupees per annum. Talking about the salary of a data scientist, a data scientist in the United States earns over $117,000 per annum, and in India, a data scientist can earn over 11 lakh rupees per annum.

Coming to the final section of this video, we'll look at the top companies hiring for data analysts, data engineers, and data scientists. So we have the first company as Google, then we have Tesla. Next we have the e-commerce giant Amazon, the internet giant Facebook, or the social media giant Facebook. We have the tech giant Oracle. We also have Verizon and Airbnb. So these are some of the top companies that hire for the three roles. If you are an aspiring data scientist who's looking out for online training and certification in data science from the best universities and industry experts, then search no more. Simply Learn's post-graduate program in data science from Caltech University in collaboration with IBM should be the right choice. For more details on this program, please use the link in the description box below.

Now let's talk about the life cycle of a data science project. Okay, the first step is the concept study. In this step, it involves understanding the business problem, asking questions, get a good understanding of the business model, meet up with all the stakeholders, understand what kind of data is available and all that is a part of the first step. So here are a few examples. We want to see what are the various specifications and then what is the end goal? What is the budget? Is there an example of this kind of a problem that has been maybe solved earlier? So all this is a part of the concept study, and another example could be a very specific one to predict the price of a 1.35-karat diamond, and there may be relevant information inputs that are available and we want to predict the price.

The next step in this process is data preparation; data gathering and data preparation, also known as data munching, or sometimes it is also known as data manipulation. So what happens here is the raw data that is available may not be usable in its current format for various reasons. So that is why in this step a data scientist would explore the data. He will take a look at some sample data. Maybe pick—if there are millions of records, pick a few thousand records and see how the data is looking. Are there any gaps? Is the structure appropriate to be fed into the system? Are there some columns which are probably not adding value? May not be required for the analysis. Very often these are like names of the customers; they will probably not add any value or much value from an analysis perspective; the structure of the data. Maybe the data is coming from multiple data sources and the structures may not be matching. What are the other problems? There may be gaps in the data. So the data—all the columns, all the cells are not filled. If you're talking about structured data, there are several blank records or blank columns. So if you use that data directly, you'll get errors or you'll get inaccurate results. So how do you either get rid of that data or how do you fill these gaps with something meaningful? So all that is a part of data munching or data manipulation. So these are some additional subtopics within that. So data integration is one of them. If there are any conflicts in the data, there may be data may be redundant. Yeah, data resident redundancy is another issue. There may be—you have, let's say, data coming from two different systems, and both of them have a customer table, for example, or customer information. So when you merge them there is a duplication issue. So how do we resolve that? So that is one. Data transformation. As I said, there will be situations where data is coming from multiple sources, and then when we merge them together they may not be matching. So we need to do some transformations to make sure everything is similar. We may have to do some data reduction. If the data size is too big, you may have to come up with ways to reduce it meaningfully without losing information. Then data cleaning. So there will be either wrong values or you null values or there are missing values. So how do you handle all of that?

A few examples of very specific stuff. So there are missing values. How do you handle missing values or null values? Here in this particular slide we are seeing three types of issues. One is missing value, then you have null value. You see the difference between the two, right? So in the missing value there is nothing—blank. Null value, it says null. Now the system cannot handle if there are null values. Similarly there is improper data. So it's supposed to be a numeric value but there is a string or a non-numeric value. So how do we clean and prepare the data so that our system can work flawlessly? So there are multiple ways, and and there is no one common way of doing this. It can vary from project to project. It can vary from what exactly is the problem you're trying to solve. It can vary from data scientist to data scientist, organization to organization. So these are like some standard practices people come up with, and and of course there will be a lot of trial and error. Somebody would have tried out something and it worked, and it'll continue to use that mechanism. So that's how we need to take care of data cleaning.

Now what are the various ways of doing, you know, if if values are missing, how do you take care of that? Now if the data is too large and um only a few records have some missing values, then it is okay to just get rid of those entire rows, for example. So if you have a million records and out of which 100 records don't have full data. So there are some missing values in about 100 records. So it's absolutely fine because it's a small percentage of the data. So you can get rid of the entire records which have missing values. But that's not a very common situation. Very often you will have multiple or at least you know a large number of data sets. For example, out of a million records you may have 50,000 records which are like having missing values. Now that's a significant amount. You cannot get rid of all those records. Your analysis will be inaccurate. So how do you handle such situations? So there are again multiple ways of doing it. One is you can probably—if a particular values are missing in a particular column, you can probably take the mean value for that particular column and fill all the missing values with the mean value so that first of all you don't get errors because of missing values and second you don't get results that are way off because these values are completely different from what is there. So that is one way. Then a few other could be either taking the median value or depending on what kind of data we are talking about. So something meaningful we will have put in there. If we are doing some machine learning activity, then obviously as a part of data preparation you need to split the data into training and test data sets. The reason being if you try to test with a data set which the system has already seen as a part of training then it will tend to give reasonably accurate results because it has already seen that data and that is not a good measure of the accuracy of the system. So typically you take the entire data set, the input data set, and split it into two parts, and again the ratio can vary from person to person, individual preferences. Some people like to split it into 50/50. Some people like it as 63.33 and 33.3. This is basically 2/3 and 1/3, and some people do it as 80/20. 80

for training and 20 for testing. So you split the data, perform the training with the 80%, and then use the remaining 20% for testing. All right.

So that is one more data preparation activity that needs to be done before you start analyzing or applying the data or putting the data through the model. Then the next step is model planning. Now this models can be statistical models. This could be machine learning models. So you need to decide what kind of models you're going to use. Again, it depends on what is the problem you're trying to solve. If it is a regression problem, you need to think of a regression algorithm and come up with a regression model. So it could be linear regression. Or if you're talking about classification, then you need to pick up an appropriate classification algorithm like logistic regression or decision tree or SVM, and then you need to train that particular model.

So that is the model building or model planning process, and the cleaned up data has to be fed into the model. And apart from cleaning, you may also have to, in order to determine what kind of model you will use, you have to perform some exploratory data analysis to understand the relationship between the various variables and u see if the data is appropriate and so on. Right? So that is the additional preparatory step that needs to be done.

So little bit of details about exploratory data analysis. So what exactly is exploratory data analysis? It is basically, as the name suggests, you're just exploring; you just received the data and you're trying to explore and uh find out what are the data types and what is the is the data clean in in each of the columns; what is the maximum minimum value. So, for example, there are out-of-the-box functionality available in tools like R. So if you just ask for a summary of the table, it will tell you for each column it will give some details as to what is the mean value, what is the maximum value and so on and so forth. So this exercise or this exploratory analysis is to get an understanding of your data, and then you can take steps to; during this process you find there are a lot of missing values; you need to take steps to fix those. You will also get an idea about what kind of model to be used and so on and so forth.

What are the various techniques used for exploratory data analysis? Typically, these would be visualization techniques like you use histograms, uh you can use box plots, you can use scatter plots. So these are very quick ways of identifying the patterns or a few of the trends of the data and so on.

And then once your data is ready, you you've decided on the model, what kind of model, what kind of algorithm you're going to use. If you're trying to do machine learning, you need to pass your 80% the training data or rather you use that training data to train your model. And the training process itself is iterative. So the training process you may have to perform multiple times, and once the training is done and you feel it is giving good accuracy then you move on to test. So you take the remaining 20% of the data. Remember we split the data into training and test. So the test data is now used to check the accuracy or how well our model is performing, and if if there are further issues let's say and model is still during testing if the accuracy is not good then you may want to retrain your model or use a different model. So this whole thing again can be iterated, but if the test process is passed or if the model passes the test then it can go into production and it will be deployed. All right.

So what are the various tools that we use for model planning? R is an excellent tool in a lot of ways. Whether you're doing regular statistical analysis or machine learning or any of these activities or in along with R studio provides a very powerful environment to do data analysis including visualization. It has a very good integrated visualization or plot mechanism which can be used for doing exploratory data analysis and then later on to do analysis detailed analysis and machine learning and so on and so forth. Then of course you can write Python programs. Python offers a rich library for performing data analysis and machine learning and so on. MATLAB is a very popular tool as well, especially during education. So this is a very easy to learn tool. So MATLAB is another uh tool that can be used. And then last but not least SAS. SAS is again very powerful. It is a preparatory tool and it has all the components that are required to perform very good statistical analysis or perform data science. So those are the various tools that would be required for or that that can be used for model building.

And uh so the next step is model building. So we have done the planning part. We said okay what is the algorithm we going to use? What kind of model we going to use? Now we need to actually train this model or build the model rather so that it can then be deployed. So what are the various uh ways or what are the various types of model building activities? So it could be let's say in this particular example that we have taken you want to find out the price of a 1.35 karat diamond. So this is let's say a linear regression problem. You have data for various carats of diamond and you use that information you pass it through a linear regression model or you create a linear regression model which can then predict your price for a 1.35 carat. So this is one example of model building and then a little bit details of how linear regression works. So linear regression is basically coming up with a relation between an independent variable and a dependent variable. So it is pretty much like coming up with the equation of a a straight line which is the best fit for the given data. So like for example here y = mx + c. So y is the dependent variable and x is the independent variable. We need to determine the values of m and c for our given data. So that is what the training process of uh this model does. At the end of the training process, you have a certain value of M and C and um that is used for predicting the values of any new data that comes. All right.

So the way it works is we use the training and the test data set to train the model and then validate whether the model is working fine or not using test data and uh if it is working fine then it is taken to the next level which is put in production. If not the model has to be retrained. If the accuracy is not good enough then the model is retrained maybe with more data or you come up with a newer model or algorithm and then repeat that process. So it is an iterative process. Once the training is completed training and test then this model is deployed and we can use this particular model to determine what is the price of a 1.35 karat diamond. Remember that was our problem statement. So now that we have the best fit for this given data, we have the price of a 1.35 karat diamond which is 10,000. So this is one example of how this whole process works.

Now how do we build the model? There are multiple ways. You can use Python for example and use libraries like pandas or numpy to build the model and implement it. This will be available as a separate tutorial, a separate video in this playlist. So stay tuned for that.

Moving on, once we have the results, the next step is to communicate this results to the appropriate stakeholders. So which is basically taking this results and preparing like a presentation or a dashboard and communicating these results to the concerned people. So finishing or getting the results of the analysis is not the last step. But you need to, as a data scientist, take this results and present it to the team that has given you this problem in the first place and explain your findings. Explain the findings of this exercise and recommend maybe what steps they need to take in order to overcome this problem or solve this problem. So that is the pretty much once that is accepted and the last step is to operationalize. So if everything is fine your data scientists presentations are accepted then they put it into practice and thereby they will be able to improve or solve the problem that they stated in step one. Okay.

So quick summary of the life cycle. You have a concept study which is basically understanding the problem asking the right questions and trying to see if there is uh enough data to solve this problem and then even maybe gather the data. Then data preparation; the raw data needs to be manipulated. You need to do data munching so that you have the data in a certain proper format to be used by the model or our analytics system. And then you need to do the model planning. What kind of a model? what algorithm you will use for a given problem and then the model building. So the exact execution of that model happens in step four and you implement and execute that model and uh put the data through the analysis in this step and then you get the results. This results are then communicated, packaged and presented and communicated to the stakeholders and once that is accepted that is operationalized. So that is the final step.

Let's begin this lesson by defining the term statistics. Statistics is a mathematical science pertaining to the collection, presentation, analysis, and interpretation of data. It's widely used to understand the complex problems of the real world and simplify them to make well-informed decisions. Several statistical principles, functions, and algorithms can be used to analyze primary data, build a statistical model, and predict the outcomes. An analysis of any situation can be done in two ways: statistical analysis or a non-statistical analysis. Statistical analysis is the science of collecting, exploring, and presenting large amounts of data to identify the patterns and trends. Statistical analysis is also called quantitative analysis. Non-statistical analysis provides generic information and includes text, sound, still images, and moving images. Non-statistical analysis is also called qualitative analysis. Although both forms of analysis provide results, statistical analysis gives more insight and a clearer picture, a feature that makes it vital for businesses.

There are two major categories of statistics: descriptive statistics and inferential statistics. Descriptive statistics helps organize data and focuses on the main characteristics of the data. It provides a summary of the data numerically or graphically. Numerical measures such as average, mode, standard deviation or SD, and correlation are used to describe the features of a data set. Suppose you want to study the height of students in a classroom. In descriptive statistics, you would record the height of every person in the classroom and then find out the maximum height, minimum height, and average height of the population. Inferential statistics generalizes the larger data set and applies probability theory to draw a conclusion. It allows you to infer population parameters based on the sample statistics and to model relationships within the data. Modeling allows you to develop mathematical equations which describe the interrelationships between two or more variables. Consider the same example of calculating the height of students in the classroom. In inferential statistics, you would categorize height as tall, medium, and small, and then take only a small sample from the population to study the height of students in the classroom.

The field of statistics touches our lives in many ways. From the daily routines in our homes to the business of making the greatest cities run, the effect of statistics are everywhere. There are various statistical terms that one should be aware of while dealing with statistics: population, sample, variable, quantitative variable, qualitative variable, discrete variable, continuous variable. A population is the group from which data is to be collected. A sample is a subset of a population. A variable is a feature that is characteristic of any member of the population differing in quality or quantity from another member. A variable differing in quantity is called a quantitative variable. For example, the weight of a person, number of people in a car. A variable differing in quality is called a qualitative variable or attribute. For example, color, the degree of damage of a car in an accident. A discrete variable is one which no value can be assumed between the two given values. For example, the number of children in a family. A continuous variable is one in which any value can be assumed between the two given values. For example, the time taken for a 100-meter run.

Typically, there are four types of statistical measures used to describe the data. They are measures of frequency, measures of central tendency, measures of spread, measures of position. Let's learn each in detail. Frequency of the data indicates the number of times a particular data value occurs in the given data set. The measures of frequency are number and percentage. Central tendency indicates whether the data values tend to accumulate in the middle of the distribution or toward the end. The measures of central tendency are mean, median, and mode. Spread describes how similar or varied the set of observed values are for a particular variable. The measures of spread are standard deviation, variance, and quartiles. The measures of spread are also called measures of dispersion. Position identifies the exact location of a particular data value in the given data set. The measures of position are percentiles, quartiles, and standard scores.

Statistical analysis system or SAS provides a list of procedures to perform descriptive statistics. They are as follows: PROC PRINT, PROC CONTENTS, PROC MEANS, PROC FREQUENCY, PROC UNIVARIATE, PROC GCHART, PROC BOXPLOT, PROC GPLOT. PROC PRINT: it prints all the variables in a SAS data set. PROC CONTENTS: it describes the structure of a data set. PROC MEANS: it provides data summarization tools to compute descriptive statistics for variables across all observations and within the groups of observations. PROC FREQUENCY: it produces one-way to n-way frequency and crosstabulation tables. Frequencies can also be an output of a SAS data set. PROC UNIVARIATE: It goes beyond what PROC MEANS does and is useful in conducting some basic statistical analyses and includes high-resolution graphical features. PROC GCHART: The GCHART procedure produces six types of charts: block charts, horizontal vertical bar charts, pie doughnut charts, and star charts. These charts graphically represent the value of a statistic calculated for one or more variables in an input SAS data set. The tread variables can be either numeric or character. PROC BOXPLOT: The BOXPLOT procedure creates side-by-side box and whisker plots of measurements organized in groups. A box and whisker plot displays the mean, quartiles, and minimum and maximum observations for a group. PROC GPLOT: GPLOT procedure creates two-dimensional graphs including simple scatter plots, overlay plots in which multiple sets of data points are displayed on one set of axes, plots against the second vertical axis, bubble plots, and logarithmic plots.

In this demo, you'll learn how to use descriptive statistics to analyze the mean from the electronic data set. Let's import the electronic data set into the SAS console. In the left plane, right-click the electronic.xlsx data set and click import data. The code to import the data generates automatically. Copy the code and paste it in the new window. The PROC MEANS procedure is used to analyze the mean of the imported data set. The keyword DATA identifies the input data set. In this demo, the input data set is electronic. The output obtained is shown on the screen. Note that the number of observations, mean, standard deviation, and maximum and minimum values of the electronic data set are obtained. This concludes the demo on how to use descriptive statistics to analyze the mean from the electronic data set.

So far you have learned about descriptive statistics. Let's now learn about inferential statistics. Hypothesis testing is an inferential statistical technique to determine whether there is enough evidence in a data sample to infer that a certain condition holds true for the entire population. To understand the characteristics of the general population, we take a random sample and analyze the properties of the sample. We then test whether or not the identified conclusions correctly represent the population as a whole. The population of hypothesis testing is to choose between two competing hypotheses about the value of a population parameter. For example, one hypothesis might claim that the wages of men and women are equal, while the other might claim that women make more than men. Hypothesis testing is formulated in terms of two hypotheses: null hypothesis which is referred to as Hnull; alternative hypothesis which is referred to as H1. The null hypothesis is assumed to be true unless there is strong evidence to the contrary. The alternative hypothesis is assumed to be true when the null hypothesis is proven false. Let's understand the null hypothesis and alternative hypothesis using a general example. Null hypothesis attempts to show that no variation exists between variables, and alternative hypothesis is any hypothesis other than the null. For example, say a pharmaceutical company has introduced a medicine in the market for a particular disease and people have been using it for a considerable period of time and it's generally considered safe. If the medicine is proved to be safe, then it is referred to as null hypothesis. To reject null hypothesis, we should prove that the medicine is unsafe. If the null hypothesis is rejected, then the alternative hypothesis is used.

Before you perform any statistical tests with variables, it's significant to recognize the nature of the variables involved. Based on the nature of the variables, it's classified into four types. They are categorical or nominal variables, ordinal variables, interval variables, and ratio variables. Nominal variables are ones which have two or more categories, and it's impossible to order the values. Examples of nominal variables include gender and blood group. Ordinal variables have values ordered logically. However, the relative distance between two data values is not clear. Examples of ordinal variables include considering the size of a coffee cup, large, medium, and small, and considering the ratings of a product, bad, good, and best. Interval variables are similar to ordinal variables except that the values are measured in a way where their differences are meaningful. With an interval scale, equal differences between scale values do have equal quantitative meaning. For this reason, an interval scale provides more quantitative information than the ordinal scale. The interval scale does not have a true zero point. A true zero point means that a value of zero on the scale represents zero quantity of the construct being assessed. Examples of interval variables include the Fahrenheit scale used to measure temperature and distance between two compartments in a train. Ratio scales are similar to interval scales in that equal differences between scale values have equal quantitative meaning. However, ratio scales also have a true zero point which give them an additional property. For example, the system of inches used with a common ruler is an example of a ratio scale. There is a true zero point because 0 inches in fact indicates a complete absence of length.

In this demo, you'll learn how to perform the hypothesis testing using SAS. In this example, let's check against the length of certain observations from a random sample. The keyword DATA identifies the input data set. The INPUT statement is used to declare the aging variable and CARDS to read data into SAS. Let's perform a t-test to check the null hypothesis. Let's assume that the null hypothesis to be that the mean days to deliver a product is 6 days. So null hypothesis equals 6. Alpha value is the probability of making an error which is 5% standard and hence alpha equals 0.05. The VAR statement names the variable to be used in the analysis. The output is shown on the screen. Note that the p-value is greater than the alpha value which is 0.05. Therefore, we fail to reject the null hypothesis. This concludes the demo on how to perform the hypothesis testing using SAS.

Let's now learn about hypothesis testing procedures. There are two types of hypothesis testing procedures: they are parametric tests and non-parametric tests. In statistical inference or hypothesis testing, the traditional tests such as t-test and ANOVA are called parametric tests. They depend on the specification of a probability distribution except for a set of free parameters. In simple words, you can say that if the population information is known completely by its parameter, then it is called a parametric test. If the population or parameter information is not known and you are still required to test the hypothesis of the population, then it's called a non-parametric test. Non-parametric tests do not require any strict distributional assumptions. There are various parametric tests. They are as follows: t-test, ANOVA, chi-squared, linear regression. Let's understand them in detail. T-test: A t-test determines if two sets of data are significantly different from each other. The t-test is used in the following situations: to test if the mean is significantly different than a hypothesized value; to test if the mean for two independent groups is significantly different; to test if the mean for two dependent or paired groups is significantly different. For example, let's say you have to find out which region spends the highest amount of money on shopping. It's impractical to ask everyone in the different regions about their shopping expenditure. In this case, you can calculate the highest shopping expenditure by collecting sample observations from each region. With the help of the t-test, you can check if the difference between the regions are significant or a statistical fluke. ANOVA:

ANOVA is a generalized version of the t-test and is used when the mean of the interval dependent variable is different from the categorical independent variable. When we want to check variance between two or more groups, we apply the ANOVA test.

For example, let's look at the same example of the t-test example. Now you want to check how much people in various regions spend every month on shopping. In this case, there are four groups, namely east, west, north, and south. With the help of the ANOVA test, you can check if the difference between the regions is significant or a statistical fluke.

Chi-square. Chi-square is a statistical test used to compare observed data with data you would expect to obtain according to a specific hypothesis. Let's understand the chi-square test through an example. You have a data set of male shoppers and female shoppers. Let's say you need to assess whether the probability of females purchasing items of $500 or more is significantly different from the probability of males purchasing items of $500 or more.

Linear regression. There are two types of linear regression: simple linear regression and multiple linear regression. Simple linear regression is used when one wants to test how well a variable predicts another variable. Multiple linear regression allows one to test how well multiple variables or independent variables predict a variable of interest. When using multiple linear regression, we additionally assume the predictor variables are independent. For example, finding the relationship between any two variables, say sales and profit, is called simple linear regression. Finding the relationship between any three variables, say sales, cost, telemarketing, is called multiple linear regression.

Some of the nonparametric tests are the Wilcoxon rank sum test and the Kruskal-Wallis H test.

Wilcoxon rank sum test. The Wilcoxon signed-rank test is a non-parametric statistical hypothesis test used to compare two related samples or matched samples to assess whether or not their population mean ranks differ. In the Wilcoxon rank sum test, you can test the null hypothesis on the basis of the ranks of the observations.

Kruskal-Wallis H test. The Kruskal-Wallis H test is a rank-based nonparametric test used to compare independent samples of equal or different sample sizes. In this test, you can test the null hypothesis on the basis of the ranks of the independent samples.

The advantages of parametric tests are as follows: Provide information about the population in terms of parameters and confidence intervals; easier to use in modeling, analyzing, and for describing data with central tendencies and data transformations; express the relationship between two or more variables; don't need to convert data into rank order to test.

The disadvantages of parametric tests are as follows: Only support normally distributed data; only applicable on variables, not attributes.

Let's now list the advantages and disadvantages of non-parametric tests. The advantages of nonparametric tests are as follows: Simple and easy to understand; do not involve population parameters and sampling theory; make fewer assumptions; provide results similar to parametric procedures.

The disadvantages of non-parametric tests are as follows: Not as efficient as parametric tests; difficult to perform operations on large samples manually.

We'll discuss the types of distribution in statistics. But before we move ahead, let's have a brief introduction on what is probability distribution. A probability distribution is a list of all of the possible outcomes of a random variable along with the corresponding probability values. And it is used in many fields, but we rarely do explain what they are. So in this video we'll discuss the three main types of probability distribution: that is, normal, binomial, and Poisson distribution. So let's move ahead.

So what is normal distribution? Normal distribution is a continuous probability density that has a probability density function which gives us a symmetrical bell curve. Now data can be distributed or spread out in different ways, but there are many cases where the data tends to be around a central value with no bias to the left or right, which means that it doesn't show any particular spikes towards the left or the right and it gets close to a normal distribution. Half of the data will fall on the left of the mean and the other half will fall on the right.

Now let's take a look at a graph which shows the height distribution in a class. As you can see, the average height is in the middle and the data to the left of the average height represents the short people and the data to the right of it represents the taller people. The y-axis shows us the likelihood of any of these heights occurring. The average height has the most distribution or it has the most number of cases in the class. And as the height decreases or increases the number of people who have that height also decreases. This kind of a distribution is called a normal distribution where the average or the mean is always the highest point and any other point after that or before that is significantly lower. The resulting data gives us a bell curve. And as you can see there is no abrupt bias or spike in the data anywhere except for the average height. So this kind of a curve is called a bell curve and it's usually seen in a normal distribution. The reason we call this a normal distribution is because the data is normally distributed with the average being the highest and all the other data points having a lower likelihood.

Now we came across two terms which are associated with normal distribution: continuous probability density and probability density function. What is continuous probability density? Continuous probability density is a probability distribution where the random variable X can take any given value because there are infinite values that X could assume. The probability of X taking on any specific value is zero. For example, let's say you have a continuous probability density for men's height. What is the probability that a man will have the exact height of 70 in? It is impossible to find this out because the probability of one man measuring exactly 70 in is very low. It is more probable that he will measure around 70.1 in or maybe 69.97 in. And it doesn't stop there. The fact is that it's impossible to exactly measure any variable that's on a continuous scale. And because of this, it's impossible to figure out the probability of one exact measurement which is occurring in a continuous probability density.

Next, we have the probability density function. It's nothing but a function or an expression which is used to define the range of values that a continuous random variable can take. An example of this would be to gauge the risk and reward of a stock. A probability density function is a statistical measure which is used to gauge the likelihood of a discrete value. A discrete variable can be measured exactly while a continuous variable can have infinite values. However, for both continuous as well as discrete variables, we can define a function which gives us the range of values within which these variables will fall. And that function is known as the probability density function.

Now let's take a look at standard deviation. What is standard deviation? Standard deviation is used to measure how the values in your data differ from one another or how spread out your data is. A standard deviation is a statistic that measures the dispersion of a data set relative to its mean. The standard deviation is calculated as the square root of variance by determining each data point's deviation relative to the mean. If the data points are further from the mean, that means that there's a higher deviation within the data set and then the data is set to be more spread out. This leads to a higher standard deviation too. Let's take an example of income in rural and urban areas. In rural areas, let's say such as farming areas, the income doesn't differ that much. More or less, everyone earns the same. Because of this, our bell curve has a very low standard deviation and it has a very narrow peak. However, in urban areas, the wealth distribution is very uneven. Some people can have very high incomes and can be earning a lot while other people can have very low incomes. Furthermore, the data distribution between these two income points is going to be more spread out because there are a lot more people living there who work in various fields and who have various incomes. Because of this, our standard deviation is more spread out and our bell curve will also have a wider peak.

Now, how can we find the standard deviation? Standard deviation is obtained by subtracting each data value from the mean and finding the squared average of these values. Let's look at how we can do this with the help of an example. These values correspond to the height of various dogs. We can find the mean by finding the average of all these values, which is nothing but adding all the values and dividing it by the total number of values. The mean that we get is 394. This means that the average height of a dog is 394 mm. To find the standard deviation, first we need to subtract the height from the mean. This will tell us how far from the mean our data points actually are. Next, we will square up all of these differences and add them up and again divide it by the total number of values that we have. This is called the variance. The variance that we get in this case is 21704. Finally, when we find the square root of this value, we will get the standard deviation. The standard deviation here is 147. The standard deviation will tell us how our data points differ from the average, and it gives us a basic value suggesting how spread out our data is from the very middle or from the mean. So if we plot these values, this value 147 will mean that a curve will have a width of 147 points around the mean.

Now what is the standard normal distribution? The standard normal distribution is a type of normal distribution that has a mean of zero and a standard deviation of one. This means that the normal distribution has its center at zero and it has intervals which increase by one. All normal distributions, like the standard normal distribution, are unimodal and symmetrically distributed with a bell-shaped curve. However, a normal distribution can take on any value as its mean and standard deviation. In the standard normal distribution, however, the mean and standard deviation are always fixed. When you standardize a normal distribution, the mean becomes zero and the standard deviation becomes one. This allows you to easily calculate the probability of certain values occurring in your distribution or to compare data sets with different mean and standard deviations. The curve shows a standard normal distribution. As you can see again the data is centered at zero. This does not mean that the data necessarily starts at zero. This means that after standardizing this point is where our mean will lie. In a standard normal distribution the standard deviation is one. So all the data points will increase or decrease in steps of one.

Let's better understand a standard normal distribution with the help of an example. Again, as you can see, the data is centered around zero, which is nothing but the mean. Let's again consider the weights of students in class 8th. The average weight here is around 50 kgs and the data increases and decreases in steps of five. The data over here in this curve is evenly distributed along these steps. This is what a standard normal distribution will look like. We already know that the mean of our data is 50. And because the data is increasing and decreasing in equal steps, we can just standardize it and take it to mean that the data is increasing and decreasing in steps of one. This is what a standard normal distribution looks like. And when you have a data which looks like this, you can always standardize it and convert it into a standard normal distribution.

Now standard normal distribution has a couple of properties which makes calculation comparatively easy. The first one is that 68% of the values fall within the first standard deviation. Which means that 68% of all data values on this curve will fall between the range of minus 1 to 1 or the first interval ranging from minus 1 to 1. The second property is that 95% of the rest of the values are within the second standard deviation or from the second negative point to the second positive point. And finally 99.7% of the values fall within the third standard deviation or from the third negative point to the third positive point. This makes calculations on standard normal distribution fairly easy. You can compare scores on different distributions with different means and standard deviations. You can normalize scores for statistical decision making using standard normal distribution. You can find the probability of observations in a distribution which fall above or below a given value. And finally, you can find the probability that a mean significantly differs from a population mean.

Now let's take a look at z-score. So what is a z-score? A z-score is used to tell us how far from the mean a data point actually is. It is calculated using the mean and standard deviation. So it can be said that the z-score is how many standard deviations below the mean our data is. Basically by using the z-score we can get an approximate location of where our data point lies on the graph with regards to the mean. Now the z-score is given by subtracting the data point from the mean and dividing it by standard deviation. This can also be written as X minus mu divided by sigma. Now any normal distribution can be standardized by converting its values into z-scores. The z-score will tell you how many standard deviations from the mean each value lies. While data points are referred to as X in a normal distribution, they are called z-scores in the z-distribution. A z-score is a standard score that will tell you how many standard deviations away from the mean an individual point will lie. A positive z-score will mean that your x value is greater than the mean and a negative z-score will mean that your x value is less than the mean. A z-score of zero will mean that your x value is equal to the mean. And again to standardize a value from a normal distribution all we have to do is convert it to a z-score by subtracting the mean from our individual value and dividing it by the standard deviation.

Now let's see how we can find the z-score from data points with the help of a solved example. Let's do a case study. In this case study we'll be taking the summary of daily travel time of a person who's commuting to and from work. All these values are in minutes and using these values we have to calculate the mean, the standard deviation and the z-score. These values are as shown. As we can see there are 13 values in total. Let's start by finding the mean. The mean is the average and it can be gotten by adding all of these values and dividing it by the total number of values. This gives us a value of 38.6. The mean tells us the average of all our data points, which means on an average, he travels for 38.6 minutes to reach work. Next, let's subtract the individual values from our mean and calculate the variance and standard deviation. The values on the left give us the values that we get after subtracting it from the mean. And the variance can be calculated by squaring all of these values, adding up all of the squared values, and dividing it by the total number of values. At the end of the day, we get a variance of 140. To calculate the standard deviation, all we have to do is take a square root of the variance, which gives us a value of 11.8. Now, the mean signifies the average of our values, and we already know this. It gives us the average time which is taken to travel. But the standard deviation will tell us the average value of how much our data points differ from the mean. It tells us the deviation within our own data and it tells us how far away on an average a point is from the mean. Now the value that we get is 11.8 which means that on an average a single data point is around 11.8 data points away from the mean. Now let's calculate the z-score. The z-score is given by subtracting individual data points from the mean and dividing it by the standard deviation. We know that we have a standard deviation of 11.8 and a mean of 38.6. Using these values, we can calculate the z-scores for individual x values. Now we know that a negative z-score means that our x value is lower than our mean. But what does the number 1.06 mean? This means that the z-score for 26 is 1.06 standard deviations away from the mean. The negative symbol here means that our x value is less than the mean. And by how less? 1.06 times the standard deviation. Now we know that the negative value of a z-score means that our x value is less than our mean. But what does the number 1.06 mean? This means that the z-score is 1.06 times the standard deviation less than the mean. The same thing can be said for the z-score of 33. It is 0.47 times the standard deviation less than the mean. The z-score of 65 is 2.23 times the standard deviation more than the mean. That means it has to be added to the mean. The reason that we know it's more than the mean is because this has a positive value. So this means that using z-scores we can know where a data point falls relative to other points on the graph. The z-score will tell us how far away from the mean a point is in steps of a standard deviation.

Basics and terminology. The first one is outcome. Whenever we do an experiment like flipping a coin or rolling a die, we get an outcome. For example, if we flip a coin, we get an outcome of heads or tails. And if we roll a die, we get an outcome of 1, 2, 3, 4, 5, or 6.

Random experiment. A random experiment is any well-defined procedure that produces an observable outcome that could not be perfectly predicted in advance. A random experiment must be well defined to eliminate any vagueness or surprise. It must produce a definite observable outcome so that you know what happened after the random experiment is run.

Random events. Consider a simple example. Let us say that we toss a coin up in the air. What can happen when it gets back? It either gives a head or a tail. These two are known as outcomes, and the occurrence of an outcome is an event. Thus, the event is the outcome of some phenomenon.

The last one is sample space. A sample space is a collection or a set of possible outcomes of a random experiment. The sample space is represented using the symbol S. The subset of all possible outcomes of an experiment is called events. And a sample space may contain a number of outcomes that depends on the experiment. If it contains a finite number of outcomes, then it is known as a discrete or finite sample space.

Now let's discuss what is random variable. A random variable is a numerical description of the outcome of a statistical experiment. A random variable that may assume only a finite number of values is set to be discrete. One that may assume any value in some interval on the real number line is set to be continuous. Let's see an example. Let X be a random variable defined as a sum of numbers when two dice are rolled. X can assume the values 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, and 12. Notice there's no one here because the sum on the two dice can never be one.

Now that we know the basics, let's move on to binomial distribution. The binomial distribution is used when there are exactly two mutually exclusive outcomes of a trial. These outcomes are appropriately labeled success and failure. The binomial distribution is used to obtain the probability of observing X successes in n number of trials with the probability of success on a single trial denoted by P. The binomial distribution assumes that P is fixed for all the trials. Here's a real-life example of a binomial distribution. Suppose you purchase a lottery ticket. Then either you are going to win the lottery or not. In other words, the outcome will be either success or failure that can be proved through binomial distribution.

There are four important conditions that need to be fulfilled for an experiment to be a binomial experiment. The first one is there should be a fixed number of entries carried out. The outcome of a given trial is only two: that is, either a success or a failure. The probability of success remains constant from trial to trial. It does not change from one trial to another. And the trials are independent. The outcome of a trial is not affected by the outcome of any other trial.

To calculate the binomial coefficient, we use the formula which is nCr * p^r * (1 - p)^(n - r) where r is the number of success in n number of trials and p is the probability of success. 1 - p denotes the probability of a failure. Now let's use this formula to solve an example. Suppose a die is tossed three times. What is the probability of no fives turning up, one five, and three fives turning up? To calculate the no fives turning up here, r

is equal to 0 and n is equal to 3. Substituting the value in the formula, we have 3C0 * (1/6)^0 * (5/6)^3, where (1/6) is the probability of success and (5/6) is the probability of failure. Calculating this equation, we'll get the value to be 0.5787.

In a similar manner, to calculate the probability of 1 out of 3 turning up, we'll replace r with one and n will be three. So P(X=1) will be equal to 3C1 * (1/6)^1 * (5/6)^2, which will come out to be 0.347. And for 3 out of 3 turning up, we substitute r equal to 3 and the formula will remain the same, and we'll get the value to be 0.046.

Now that we are done with the concepts of binomial probability distribution, here's a problem for you to solve. Post your answers in the comment section and let us know.

A Poisson distribution is a probability distribution used in statistics to show how many times an event is likely to happen over a given period of time. To put it another way, it's a count distribution. Poisson distributions are frequently used to comprehend independent events at a constant rate over a given interval of time. The Poisson distribution was developed by French mathematician Simon Denis Poisson in 1837. A Poisson distribution is used in cases where the chances of any individual event being a success is very small. The number of defective pencils per box of 6,000 pencils, the number of plane crashes in India in one year, or the number of printing mistakes in each page of a book. All of these examples can have use of Poisson distribution. The Poisson distribution can be used to calculate how likely it is that something will happen X number of times. A random variable X has a Poisson distribution with parameter lambda. And the formula for that is e^-λ * λ^x / x!, where x can be the number of times the event is happening. The value of e is taken as 2.7182.

Let's discuss some applications of Poisson distribution. If you want to calculate the number of deaths per day or week due to a rare disease in a hospital, you can use a Poisson distribution. In a similar manner, the count of bacteria per cc in blood or the number of computers infected with a virus per week. The number of mishandled baggage per thousand passengers can also have an application for Poisson distribution.

Let's discuss one example to see how we can calculate the Poisson distribution. Suppose on average a cancer kills five people each year in India. What is the probability that one person is killed this year? We'll assume all these events are independent random events. So by the formula, we have x = 1 because we have to calculate the probability of 1 person that is killed this year. So P(X=1) will be equal to e^-5 * 5^1 / 1!, which will come out to be 0.033, which will be near to 3.3%. So the probability that only one person is killed this year due to cancer is 3.3%.

If you are an aspiring data scientist who is looking out for online training and certification in data science from the best universities and industry experts, then search no more; simply learn's post-graduate program in data science from Caltech University in collaboration with IBM should be the right choice. For more details on this program, please use the link in the description box below.

Hello everyone, welcome to another session by simply learn. Today we are going to discuss Bayes' theorem, an important subtopic that comes under probability theory. We'll start this video by talking about probability and conditional probability. After that, we'll move on to Bayes' theorem and understand its formula and a real-life example where Bayes' theorem can be used. So let's get started.

What is probability? Probability is a branch of mathematics concerning numerical descriptions of how likely an event is to occur or how likely it is that a proposition is true. The probability of an event is a number between 0 and 1. Well, roughly speaking, zero indicates the impossibility of the event and one indicates certainty. The higher the probability of an event, the more likely it is that the event will occur. Let's look at an example. A simple example is the tossing of a fair unbiased coin. Since the coin is fair, the outcome that is heads and tails are both equally probable. The probability of heads equals the probability of tails. And since no other outcomes are possible, the probability of either heads or tails can be set to be 1/2, which is also 50%. The probability of an event can be calculated by the number of ways it can happen divided by the total number of outcomes.

Now that we know about probability, let's see if you can answer this question: What is the probability of drawing a jack and a queen consecutively from a deck of 52 cards without replacement? Here are your options. Post your answers in the comment section and let us know.

Now let's move on to conditional probability. Let A and B be the two events associated with a random experiment. Then the probability of A's occurrence under the condition that B has already occurred and probability of B is not equal to zero is called the conditional probability. It is denoted by P(A|B). Thus we can say that P(A|B) is equal to P(A ∩ B) / P(B), where P(A|B) is the probability of occurrence of A given that B has already occurred and P(B) is the probability of occurrence of B. To know more about conditional probability, you can check our previous video which is specifically on conditional probability.

Now let's move on to Bayes' theorem. Bayes' theorem is a mathematical formula for calculating conditional probability in probability and statistics. In other words, it is used to figure out how likely an event is associated based on its proximity to another. Bayes' law or Bayes' rule are the other names of this theorem. The formula for Bayes' theorem can be written in a variety of ways. The most common version is P(A|B) = P(B|A) * P(A) / P(B). Where P(A|B) is the conditional probability of event A occurring given that B is true, and P(A) and P(B) are the probabilities of A and B occurring independently of one another.

Let's solve a problem using Bayes' theorem to understand it better. There is a cricket match tomorrow, and in recent years it has rained only 5 days each year. Unfortunately, the meteorologist has predicted rain for tomorrow. Now when it rains, the meteorologist correctly forecasts rain 90% of the time, and when it doesn't rain, he incorrectly forecasts rain 10% of the time. Let's calculate what is the probability that it will rain on the match day. So the two sample spaces here are the events that it rains and it does not rain. Additionally, a third event is also there that the meteorologist predicts the rain. So the notation for these events appears below. Event A1 = it rains on the match day. Event A2 = it does not rain on the match day, and event B = the meteorologist predicting the rain. Now in terms of probability, we know the following: P(A1) = 5/365, that it rains 5 days in a year, which will come out to be 0.0136. P(A2) = 360/365, that is no rain for 360 days in a year, which will come out to be 0.986. P(B|A1) = 0.9. This signifies when it rains, the meteorologist predicts the rain 90% of the time. In a similar manner, P(B|A2) = 0.1, that it does not rain; the meteorologist predicts the rain 10% of the time. Combining all this, we can calculate P(A1|B), that is the probability it will rain on the given match day given a forecast of rain by the meteorologist. The answer can be determined using Bayes' theorem as shown below. So here's the formula of Bayes' theorem, and putting all the values that we have calculated in the previous slide. The probability that it will rain on the match day given a forecast of rain by the meteorologist will come out to be 0.111, which will be equal to 11.11%. So there's an 11% chance that it will rain on the match day given that the meteorologist has predicted the rain. I hope this example is clear to you.

It's a weekend, and John decided to watch the latest movie recommended by Netflix at his friend's place. Before heading out, he asked Siri about the weather and realized it would rain. So, he decided to take his Tesla for the long journey and switched to autopilot on the highway. After coming home from the eventful day, he started wondering how technology has made his life easy. He did some research on the internet and found out that Netflix, Siri, and Tesla are all using AI. So, what is AI? AI, or artificial intelligence, is nothing but making computer-based machines think and act like humans. Artificial intelligence is not a new term. John McCarthy, a computer scientist, coined the term artificial intelligence back in 1956, but it took time to evolve as it demanded heavy computing power. Artificial intelligence is not confined to just movie recommendations and virtual assistants. Broadly classifying, there are three types of AI: Artificial narrow intelligence, also called weak AI, is the stage where machines can perform a specific task. Netflix, Siri, chatbots, facial recommendation systems are all examples of artificial narrow intelligence. Next up, we have artificial general intelligence, referred to as an intelligent agent's capacity to comprehend or pick up any intellectual skill that a human can. We are halfway into successfully implementing this space. IBM's Watson supercomputer and GPT-3 fall under this category. And lastly, artificial super intelligence. It is the stage where machines surpass human intelligence. You might have seen this in movies and imagined how the world would be if machines occupied it.

Fascinated by this, John did more research and found out that machine learning, deep learning, and natural language processing are all connected with artificial intelligence. Machine learning, a subset of AI, is the process of automating and enhancing how computers learn from their experiences without human help. Machine learning can be used in email spam detection, medical diagnosis, etc. Deep learning can be considered a subset of machine learning. It is a field that is based on learning and improving on its own by examining computer algorithms. While machine learning uses simpler concepts, deep learning works with artificial neural networks which are designed to imitate the human brain. This technology can be applied in face recognition, speech recognition, and many more applications. Natural language processing, popularly known as NLP, can be defined as the ability of machines to learn human language and translate it. Chatbots fall under this category. Artificial intelligence is advancing in every crucial field like healthcare, education, robotics, banking, e-commerce, and the list goes on. Like in healthcare, AI is used to identify diseases, helping healthcare service providers and their patients make better treatment and lifestyle decisions. Coming to the education sector, AI is helping teachers automate grading, organizing, and facilitating parent-guardian conversations. In robotics, AI-powered robots employ real-time updates to detect obstructions in their path and instantaneously design their routes. Artificial intelligence provides advanced data analytics that is transforming banking by reducing fraud and enhancing compliance. With this growing demand for AI, more and more industries are looking for AI engineers who can help them develop intelligent systems and offer them lucrative salaries going north of $120,000. The future of AI looks promising, with the AI market expected to reach $190 billion by 2025.

We know humans learn from their past experiences, and machines follow instructions given by humans. But what if humans can train the machines to learn from their past data and do what humans can do and much faster? Well, that's called machine learning. But it's a lot more than just learning; it's also about understanding and reasoning. So today we will learn about the basics of machine learning.

So that's Paul. He loves listening to new songs. He either likes them or dislikes them. Paul decides this on the basis of the song's tempo, genre, intensity, and the gender of voice. For simplicity, let's just use tempo and intensity for now. So here tempo is on the x-axis ranging from relaxed to fast, whereas intensity is on the y-axis ranging from light to soaring. We see that Paul likes the song with fast tempo and soaring intensity while he dislikes the song with relaxed tempo and light intensity. So now we know Paul's choices. Let's say Paul listens to a new song. Let's name it as song A. Song A has fast tempo and a soaring intensity. So it lies somewhere here. Looking at the data, can you guess whether Paul will like the song or not? Correct. So Paul likes this song. By looking at Paul's past choices, we were able to classify the unknown song very easily, right? Let's say now Paul listens to a new song. Let's label it as song B. So song B lies somewhere here with medium tempo and medium intensity. Neither relaxed nor fast, neither light nor soaring. Now, can you guess whether Paul likes it or not? Not able to guess whether Paul will like it or dislike it. Are the choices unclear? Correct. We could easily classify song A. But when the choice became complicated as in the case of song B, yes, and that's where machine learning comes in. Let's see how. In the same example for song B, if we draw a circle around song B, we see that there are four votes for like whereas one vote for dislike. If we go for the majority votes, we can say that Paul will definitely like the song. That's all. This was a basic machine learning algorithm also. It's called K-nearest neighbors. So this is just a small example of one of the many machine learning algorithms. Quite easy, right? Believe me, it is. But what happens when the choices become complicated as in the case of song B? That's when machine learning comes in. It learns the data, builds the prediction model, and when the new data point comes in, it can easily predict for it. The more the data, the better the model, the higher will be the accuracy. There are many ways in which the machine learns. It could be either supervised learning, unsupervised learning, or reinforcement learning. Let's first quickly understand supervised learning. Suppose your friend gives you 1 million coins of three different currencies, say 1 rupee, 1 euro, and 1 dirham. Each coin has different weights. For example, a coin of one rupee weighs 3 g, 1 euro weighs 7 g, and 1 dirham weighs 4 g. Your model will predict the currency of the coin. Here your weight becomes the feature of coins. Currency becomes the label. When you feed this data to the machine learning model, it learns which feature with which label. For example, it will learn that if a coin is of 3 g, it will be a 1 rupee coin. Let's give a new coin to the machine. On the basis of the weight of the new coin, your model will predict the currency. Hence, supervised learning uses labeled data to train the model. Here, the machine knew the features of the object and also the labels associated with those features. On this note, let's move to unsupervised learning and see the difference. Suppose you have a cricket data set of various players with their respective scores and wickets taken. When we feed this data set to the machine, the machine identifies the pattern of player performance. So, it plots this data with the respective wickets on the x-axis while runs on the y-axis. While looking at the data, you'll clearly see that there are two clusters. One cluster is the players who scored high runs and took less wickets, while the other cluster is of the players who scored less runs but took many wickets. So here we interpret these two clusters as batsmen and bowlers. The important point to note here is that there were no labels of batsmen and bowlers. Hence, the learning with unlabeled data is unsupervised learning. So we saw supervised learning where the data was labeled and unsupervised learning where the data was unlabeled. And then there is reinforcement learning, which is a reward-based learning, or we can say that it works on the principle of feedback. Here, let's say you provide the system with an image of a dog and ask it to identify it. The system identifies it as a cat. So you give negative feedback to the machine saying that it's a dog's image. The machine will learn from the feedback, and finally, if it comes across any other image of a dog, it'll be able to classify it correctly. That is reinforcement learning. To generalize the machine learning model, let's see a flowchart. Input is given to a machine learning model, which then gives the output according to the algorithm applied. If it's right, we take the output as a final result. Else, we provide feedback to the training model and ask it to predict until it learns. I hope you've understood supervised and unsupervised learning. So let's have a quick quiz. You have to determine whether the given scenarios use supervised or unsupervised learning. Simple, right? Scenario one: Facebook recognizes your friend in a picture from an album of tagged photographs. Scenario two: Netflix recommends new movies based on someone's past movie choices. Scenario three: Analyzing bank data for suspicious transactions and flagging the fraud transactions. Think wisely and comment below your answers.

Moving on, don't you sometimes wonder how is machine learning possible in today's era? Well, that's because today we have humongous data available. Everybody's online, either making a transaction or just surfing the internet, and that's generating a huge amount of data every minute, and that data, my friend, is the key to analysis. Also, the memory handling capabilities of computers have largely increased, which helps them to process such a huge amount of data at hand without any delay. And yes, computers now have great computational powers. So there are a lot of applications of machine learning out there. To name a few, machine learning is used in healthcare where diagnostics are predicted for doctor's review. The sentiment analysis that the tech giants are doing on social media is another interesting application of machine learning. Fraud detection in the finance sector and also to predict customer churn in the e-commerce sector. While booking a cab, you must have encountered surge pricing often where it says the fare of your trip has been updated. Continue booking. Yes, please. I'm getting late for office. Well, that's an interesting machine learning model which is used by global taxi giant Uber and others where they have differential pricing in real-time based on demand, the number of cars available, bad weather, rush hour, etc. So they use the surge pricing model to ensure that those who need a cab can get one. Also, it uses predictive modeling to predict where the demand will be high with a goal that drivers can take care of the demand and surge pricing can be minimized.

Great. Hey Siri, can you remind me to book a cab at 6 p.m. today? Okay, I'll remind you. Thanks. No problem.

If you are an aspiring data scientist who is looking out for online training and certification in data science from the best universities and industry experts, then search no more. Simply learn's post-graduate program in data science from Caltech University in collaboration with IBM should be the right choice. For more details on this program, please use the link in the description box below.

Let's dive in a little deeper and see how machine learning works. Let's say you provide a system with input data that carries the photos of various kinds of fruits. Now you want the system to figure out what are the different fruits and group them accordingly. So what does the system do? It analyzes the input data. Then it tries to find patterns, patterns like shapes, size, and color. Based on these patterns, the system will try to predict the different types of fruit and segregate them. Finally, it keeps track of all such decisions it took in the process to make sure it's learning. The next time you ask the same system to predict and segregate the different types of fruits, it won't have to go through the entire process again. That's how machine learning works. Now let's look into the types of machine learning. Machine learning is primarily of three types. First one is supervised machine learning. As the name suggests, you have to supervise your machine learning while you train it to work on its own. It requires labeled training data. Next up is unsupervised learning wherein there will be training data, but it won't be labeled. Finally, there's reinforcement learning wherein the system learns on its own. Let's talk about all these types in detail. Let's try to understand how supervised learning works. Look at the pictures very, very carefully. The monitor depicts the model or the system that we are going to train. This is how the training is done. We provide a data set that contains pictures of a kind of fruit, say an apple. Then we provide another data set which lets the model know that these pictures were that of a fruit called apple. This ends the training phase. Now what we will do is we provide a new set of data which only contains pictures of apples. Now here comes the fun part. The system can actually tell you what fruit it is, and it will remember this and apply this knowledge in the future as well. That's how supervised learning works. You are training the model to do a certain kind of operation on its own. This kind of

A model is generally used to filtering spam mails from your email accounts as well. Yes, surprise, aren't you?

So, let's move on to unsupervised learning. Now, let's say we have a data set which is cluttered. In this case, we have a collection of pictures of different fruits. We feed this data to the model, and the model analyzes the data to figure out patterns in it. In the end, it categorizes the photos into three types, as you can see in the image, based on their similarities. So, you provide the data to the system and let the system do the rest of the work. Simple, isn't it? This kind of a model is used by Flipkart to figure out the products that are well suited for you. Honestly speaking, this is my favorite type of machine learning out of all the three. And this type has been widely shown in most of the sci-fi movies lately.

Let's find out how it works. Imagine a newborn baby. You put a burning candle in front of the baby. The baby does not know that if it touches the flame, its fingers might get burned. So, it does that anyway and gets hurt. The next time you put that candle in front of the baby, it will remember what happened the last time and would not repeat what it did. That's exactly how reinforcement learning works. We provide the machine with a data set wherein we ask it to identify a particular kind of a fruit. In this case, an apple. So what it does as a response, it tells us that it's a mango. But as we all know, it's a completely wrong answer. So as a feedback, we tell the system that it's wrong. It's not a mango; it's an apple. What it does, it learns from the feedback and keeps that in mind. When the next time when we ask a same question, it gives us the right answer. It is able to tell us that it's actually an apple. That is a reinforced response. So that's how reinforcement learning works. It learns from its mistakes and experiences. This model is used in games like Prince of Persia, or Assassin's Creed, or FIFA, wherein the level of difficulty increases as you get better with the games.

Just to make it more clear for you, let's look at a comparison between supervised and unsupervised learning. Firstly, the data involved in case of supervised learning is labeled. As we mentioned in the examples previously, we provide the system with a photo of an apple and let the system know that this is actually an apple. That is called label data. So, the system learns from the label data and makes future predictions. Now unsupervised learning does not require any kind of label data because its work is to look for patterns in the input data and organize it. The next point is that you get a feedback in case of supervised learning. That is, once you get the output, the system tends to remember that and uses it for the next operation. That does not happen for unsupervised learning. And the last point is that supervised learning is mostly used to predict data, whereas unsupervised learning is used to find out hidden patterns or structures in data. I think this would have made a lot of things clear for you regarding supervised and unsupervised learning.

Now let's talk about a question that everyone needs to answer before building a machine learning model: What kind of a machine learning solution should we use? Yes, you should be very careful with selecting the right kind of solution for your model because if you don't, you might end up losing a lot of time, energy, and processing cost. I won't be naming the actual solutions because you guys aren't familiar with them yet. So, we will be looking at it based on supervised, unsupervised, and reinforcement learning. So, let's look into the factors that might help us select the right kind of machine learning solution.

First factor is the problem statement; it describes the kind of model you will be building, or as the name suggests, it tells you what the problem is. For example, let's say the problem is to predict the future stock market prices. So for anyone who is new to machine learning would have trouble figuring out the right solution. But with time and practice, you will understand that for a problem statement like this, a solution based on supervised learning would work the best for obvious reasons. Then comes the size, quality, and nature of the data. If the data is cluttered, you go for unsupervised. If the data is very large and categorical, we normally go for supervised learning solutions. Finally, we choose a solution based on their complexity. As for the problem statement wherein we predict the stock market prices, it can also be solved by using reinforcement learning. But that would be very, very difficult and time-consuming unlike supervised learning.

Algorithms are not types of machine learning. In the most simplest language, they are methods of solving a particular problem. So the first kind of method is classification, which falls under supervised learning. Classification is used when the output you are looking for is a yes or a no, or in the form A or B, or true or false. Like if a shopkeeper wants to predict if a particular customer will come back to his shop or not, he will use a classification algorithm. The algorithms that fall under classification are decision tree, naive Bayes, random forest, logistic regression, and KNN. The next kind is regression. This kind of a method is used when the predicted data is numerical in nature. Like if the shopkeeper wants to predict the price of a product based on its demand, it would go for regression. The last method is clustering. Clustering is a kind of unsupervised learning. Again, it is used when the data needs to be organized. Most of the recommendation systems used by Flipkart, Amazon, etc. make use of clustering. Another major application of it is in search engines. The search engines study your old search history to figure out your preferences and provide you the best search results. One of the algorithms that fall under clustering is K means.

Now that we know the various algorithms, let's look into four key algorithms that are used widely. We will understand them with very simple examples. The four algorithms that we will try to understand are K nearest neighbor, linear regression, decision tree, and naive Bayes.

Let's start with our first machine learning solution: K nearest neighbor. K nearest neighbor is again a kind of a classification algorithm. As you can see on the screen, the similar data points form clusters: the blue one, the red one, and the green one. There are three different clusters. Now, if we get a new and unknown data point, it is classified based on the cluster closest to it or the most similar to it. K in KNN is the number of nearest neighboring data points we wish to compare the unknown data with. Let's make it clear with an example. Let's say we have three clusters in a cost to durability graph. First cluster is of footballs. The second one is of tennis balls, and the third one is of basketballs. From the graph we can say that the cost of footballs is high and the durability is less. The cost of tennis balls is very less, but the durability is high, and the cost of basketballs is as high as the durability. Now let's say we have an unknown data point. We have a black spot which can be one kind of the balls, but we don't know what kind it is. So what we'll do, we'll try to classify this using KNN. So if we take K is equal to 5, we draw a circle keeping the unknown data point as the center, and we make sure that we have five balls inside that circle. In this case, we have a football, a basketball, and three tennis balls. Now since we have the highest number of tennis balls inside the circle, the classified ball would be a tennis ball. So that's how K nearest neighbor classification is done.

Linear regression is again a type of supervised learning algorithm. This algorithm is used to establish a linear relationship between variables, one of which would be dependent and the other one would be independent. Like if we want to predict the weight of a person based on his height, weight would be the dependent variable and height would be independent. Let's have a look at it through an example. Let's say we have a graph here showing a relationship between height and weight of a person. Let's put the y-axis as age and the x-axis as weight. So the green dots are the various data points. These green dots are the data points, and D is the mean squared error. That is, the perpendicular distances from the line to the data points are the error values. This error tells us how much the predicted values vary from the original value. Let's ignore this blue line for a while. So let's say if this is our regression line, you can see the distance from all the data points from this line is very high. So if we take this line as a regression line, the error in the prediction will be too high. So in this case, the model will not be able to give us a good prediction. Let's say we draw another regression line here like this. Even in this case, you can see that the perpendicular distance of the data points from the line is very high. So the error value will still come as high as the last one. So this model will also not be able to give us a good prediction. So what to do? So finally, we draw a line, which is this blue line. So here we can see that the distance of the data points from the line is very less relative to the other two lines we drew. So the value of D for this line will be very less. So in this case, if we take any value on the x-axis, the corresponding value on the y-axis will be our prediction. And given the fact that the D is very low, our prediction should be good also. This is how regression works. We draw a line, a regression line, that is in such a way that the value of D is the least, eventually giving us good predictions.

This algorithm, that is decision tree, is a kind of an algorithm you can very strongly relate to. It uses a kind of a branching method to realize the problem and make decisions based on the conditions. Let's take this graph as an example. Imagine yourself sitting at home getting bored. You feel like going for a swim. What you do is you check if it's sunny outside. So that's your first condition. If the answer to that condition is yes, you go for a swim. If it's not sunny, then the next question you would ask yourself is if it's raining outside. So that's condition number two. If it's actually raining, you cancel the plan and stay indoors. If it's not raining, then you would probably go outside and have a walk. So that's the final node. That's how decision tree algorithm works. You probably use this every day. It realizes a problem and then takes the decisions based on the answers to every condition.

Naive Bayes algorithm is mostly used in cases where a prediction needs to be done on a very large data set. It makes use of conditional probability. Conditional probability is the probability of an event, say A, happening given that another event B has already happened. This algorithm is most commonly used in filtering spam mails in your email account. Let's say you receive a mail. The model goes through your old spam mail records. Then it uses Bayes' theorem to predict if the present mail is a spam mail or not. So P(C|A) is the probability of event C occurring when A has already occurred. P(A|C) is the probability of event A occurring when C has already occurred. And P(C) is the probability of event C occurring, and P(A) is the probability of event A occurring. Let's try to understand naive Bayes with a better example. Naive Bayes can be used to determine on which days to play cricket based on the probabilities of a day being rainy, windy, or sunny. The model tells us if a match is possible. If we consider all the weather conditions to be event A for us and the probability of a match being possible event C. So the model applies the probabilities of event A and C into the Bayes' theorem and predicts if a game of cricket is possible on a particular day or not. In this case, if the probability of C|A is more than 0.5, we can be able to play a game of cricket. If it's less than 0.5, we won't be able to do that. That's how naive Bayes algorithm works.

We're going to cover reinforcement learning today and what's in it for you. We'll start with why reinforcement learning. We'll look at what is reinforcement learning. We'll see what the different kinds of learning strategies are that are being used today in computer models under supervised versus unsupervised versus reinforcement. We'll cover important terms specific to reinforcement learning. We'll talk about Markov's decision process, and we'll take a look at a reinforcement learning example. Well, we'll teach a tic-tac-toe how to play.

Why reinforcement learning? Training a machine learning model requires a lot of data, which might not always be available to us. Further, the data provided might not be reliable. Learning from a small subset of actions will not help expand the vast realm of solutions that may work for a particular problem. And you can see here we have the robot learning to walk. Um, very complicated setup when you're learning how to walk. And you'll start asking questions like, if I'm taking one step forward and left, what happens if I pick up a 50 lb object? How does that change how a robot would walk? These things are very difficult to program because there's no actual information on it until the it's actually tried out. Learning from a small subset of actions will not help expand the vast realm of solutions that may work for a particular problem. And we'll see here it learned how to walk. This is going to slow the growth that technology is capable of. Machines need to learn to perform actions by themselves and not just learn off humans. And you see the objective: climb a mountain. Real interesting point here is that as human beings, we can go into a very unknown environment and we can adjust for it and kind of explore and play with it. Most of the models, the non-reinforcement models in computer machine learning aren't able to do that very well. Uh, there's a couple of them that can be used or integrated. See how it goes is what we're talking about with reinforcement learning.

So what is reinforcement learning? Reinforcement learning is a subbranch of machine learning that trains a model to return an optimum solution for a problem by taking a sequence of decisions by itself. Consider a robot learning to go from one place to another. The robot is given a scenario; it must arrive at a solution by itself. The robot can take different paths to reach the destination. It will know the best path by the time taken on each path. It might even come up with a unique solution all by itself. And that's really important as we're looking for unique solutions. Uh, we want the best solution, but you can't find it unless you try it. So we're looking at our different systems, our different models. We have supervised versus unsupervised versus reinforcement learning. And with the supervised learning, that is probably the most controlled environment. Uh, we have a lot of different supervised learning models, whether it's linear regression, neural networks, um, there's all kinds of things in between, decision trees. The data provided is labeled data with output values specified. And this is important because when we talk about supervised learning, you already know the answer for all this information. You already know the picture has a motorcycle in it. So you're supervised learning. You already know that um the outcome for tomorrow for, you know, going back a week. You're looking at stock. You can already have like the graph of what the next day looks like. So you have an answer for it. And you have labeled data which is used. You have an external supervision and solves problems by mapping labeled input to known output. So very controlled.

Unsupervised learning, and unsupervised learning is really interesting because it's now taking part in many other models. They start with an; you can actually insert an unsupervised learning model um in almost either supervised or reinforcement learning as part of the system, which is really cool. Uh, data provided is unlabeled data. The outputs are not specified. The machine makes its own predictions used to solve association with clustering problems. Unlabeled data is used. No supervision. Solves problems by understanding patterns and discovering output. Uh, so you can look at this and you can think um some of these things go with each other. They belong together. So it's looking for what connects in different ways. And there's a lot of different algorithms that look at this. Um, when you start getting into those, there's some really cool images that come up of what unsupervised learning is. How it can pick out, say, uh the area of a donut. One model will see the area of the donut, and the other one will divide it into three sections based on its location versus what's next to it. So there's a lot of stuff that goes in with unsupervised learning.

And then we're looking at reinforcement learning. Probably the biggest industry in today's market uh in machine learning or growing market. It's very in its very infant stage uh as far as how it works and what it's going to be capable of. The machine learns from its environment using rewards and errors used to solve reward-based problems. No predefined data is used. No supervision follows trail and error problem-solving approach. Uh, so again we have a random; at first you start with a random: I try this; it works, and this is my reward. Doesn't work very well maybe or maybe doesn't even get you where you're trying to get it to do, and you get your reward back, and then it looks at that and says, well, let's try something else, and it starts to play with these different things finding the best route.

So let's take a look at important terms in today's reinforcement model, and this has become pretty standardized over the last few years, so these are really good to know. We have the agent; agent is the model that is being trained via reinforcement learning. So this is your actual u entity that has however you're doing it, whether you're using a neural network or a Q table or whatever combination thereof. This is the actual agent that you're using. This is the model, and you have your environment. Uh, the training situation that the model must optimize to is called its environment. Uh, and you can see here I guess we have a robot who's trying to get a chest full of gyms or whatever. And that's the output. And then you have your action. This is all possible steps that can be taken by the model, and it picks one action, and you can see here it's picked three different uh routes to get to the chest of diamonds and gems. We have a state: the current position condition returned by the model. And you could look at this uh if you're playing like a video game. This is the screen you're looking at. Uh, so when we go back here, the environment is a whole game board. So, if you're playing one of those Mobius games, you might have the whole game board going on. Uh, but then you have your current position. Where are you on that game board? What's around that? What's around you? Um, if you were talking about a robot, the environment might be moving around the yard, where it is in the yard, and what it can see, what input it has in that location. That would be the current position condition returned by the model. And then the reward uh to help the model move in the right direction. It is rewarded. Points are given to it to appraise some kind of action. So yeah, you did good, or if uh didn't do as good, trying to maximize the reward and have the best reward possible. And then policy. Policy determines how an agent will behave at any time. It acts as a mapping between action and present state. This is part of the model. What what is your action that you're you're going to take? What's the policy you're using to have an output from your agent? One of the reasons they separate a policy as its own entity is that you usually have a prediction um of a different options, and then the policy: well, how am I going to pick the best based on those predictions? I'm going to guess at different options, and we'll actually weigh those options in and find the best option we think will work. Uh, so it's a little tricky, but the policy thing is actually pretty cool how it works.

Let's go ahead and take a look at a reinforcement learning example. And just in looking at this, we're going to take a look; consider what a dog um that we want to train. Uh, so the dog would be like the agent. So you have your your puppy or whatever. Uh, and then your environment is going to be the whole house or whatever it is where you're training them. And then you have an action. We want to teach the dog to

fetch. So, action equals fetching. Uh, and then we have a little biscuit. So we can get the dog to perform various actions by offering incentives such as a dog biscuit as a reward. The dog will follow a policy to maximize this reward and hence will follow every command and might even learn new actions like begging by itself.

Uh, so you have b, you know, so we start off with fetching. It goes, "Oh, I get a biscuit for that." It tries something else and you get a handshake or begging or something like that, and it goes, "Oh, this is also reward-based," and so it kind of explores things to find out what will bring it a biscuit, and that's very much like how a reinforced model goes. Is it uh looks for different rewards? How do I find, can I try different things and find a reward that works? The dog also will want to run around and play and explore its environment. Uh, this quality of model is called exploration, so there's a little randomness going going on in exploration and explores new parts of the house. Climbing on the sofa doesn't get a reward. In fact, it usually gets kicked off the sofa.

So, let's talk a little bit about Markov's decision process. Uh, Markov's decision process is a reinforcement learning policy used to map a current state to an action where the agent continuously interacts with the environment to produce new solutions and receive rewards. And you'll see here's all of our different uh uh vocabulary we just went over. We have a reward, our state, our agent, our environment, interaction. And so even though the environment kind of contains everything um that you you really when you're actually writing the program, your environment is going to put out a reward and state that goes into the agent. Uh, the agent then looks at this uh state or it looks at the reward usually um first and it says, "Okay, I got rewarded for whatever I just did or I didn't get rewarded," and then it looks at the state and then it comes back and if you remember from policy, the policy comes in um and then we have a reward. The policy is that part that's connected at the bottom. And so it looks at that policy and it says, "Hey, what's a good action that will probably be similar to what I did?" Or um uh sometimes they're completely random, but what's a good action that's going to bring me a different reward? So, taking the time to just understand these different pieces as they go is pretty important in most of the models today.

Um, and so a lot of them actually have templates based on this that you can pull in and start using. Um, pretty straightforward as far as once you start seeing how it works, uh, you can see your environment sends it says, "Hey, this is the agent did this. If you're a character in a game, this happened," and it shoots out a reward in a state. The agent looks at the reward, looks at the new state, and then takes a little guess and says, "I'm going to try this action." And then that action goes back into the environment; it affects the environment. The environment then changes depending on what the action was, and then it has a new state and a new reward that goes back to the agent.

So, in the diagram shown, we need to find the shortest path between node A and D. Each path has a reward associated with it, and the path with a maximum reward is what we want to choose. The nodes A, B, C, D denote the nodes to travel from node uh A to B is an action. Reward is the cost of each path, and policy is each path taken. And you can see here A can go uh to B or A can go to C right off the bat or it can go right to D. And if you explored all three of these uh you would find that A going to D was a zero reward. Um, A going to C and D would generate a different reward. Or you could go A, C, B, D. There's a lot of options here.

Um, and so when we start looking at this diagram, you start to realize that even though uh today's reinforced learning models do really good at um finding an answer, they end up trying almost all the different directions you see. And so they take up a lot of work uh or a lot of processing time for reinforcement learning. They're right now in their infant stage, and they're really good at solving simple problems. And we'll take a look at one of those in just a minute in a tic-tac-toe game. Uh, but you can see here uh once it's gone through these and it's explored, it's going to find that A, C, D is the best reward. It gets a full 30 points for it.

So let's go ahead and take a look at a reinforcement learning demo. Uh, in this demo, we're going to use reinforcement learning to make a tic-tac-toe game. You'll be playing this game against the machine learning model. And we'll go ahead and we're doing it in Python. So, let's go ahead and go through um I always uh not always actually have a lot of Python tools. Let's go through um Anaconda, which will open up a Jupyter notebook. Seems like a lot of steps, but it's worth it to keep all my stuff separate, and it's also has a nice display when you're in the Jupyter notebook for doing Python.

So, here's our Anaconda Navigator. I open up the notebook, which is going to take me to a web page. And I've gone in here and created a new uh Python folder. In this case, I've already done it and enabled it. Change the name to tic-tac-toe. Uh, and then for this example, uh, we're going to go ahead and import a couple things. We're going to, um, import numpy as np. We'll go ahead and import pickle. Numpy, of course, is our number array. And, uh, pickle is just a nice way sometimes for storing, uh, different information, uh, different states that we're going to go through on here. Uh, and so we're going to create a class called state. I'm going to start with that. And there's a lot of uh lines of code to this uh class that we're going to put in here. Don't let that scare you too much. There's not as much here. Um, it looks like there's going to be a lot here, but there really is just a lot of setup going on in the in our class state. And so we have up here, we're going to initialize it. Um, we have our board. Um, it's a tic-tac-toe board, so we're only dealing with nine spots on the board. Uh, we have player one, player two, uh, is end. We're going to create a board hash. U, we'll look at that in just a minute. We're just going to store some information in there. Symbol of player equals one. Um, so there's a few things going on as far as the initialization. Uh, then something simple. We're just going to get the hash um of the board. You get the information from the board on there, which is uh columns and rows. We want to know when a winner occurs. Uh, so if you get three in a row, that's what this whole section here is for. Uh, let me go ahead and scroll up a little bit. And you can get a copy of this code if you send a note over to SimplyLearn. We'll send you over um this particular file and you can play with it yourself and see how it's put together. I don't want to spend a huge amount of time on this uh because this is just some real general Python coding. Uh, but you can see here we're just going through all the rows and you add them together and if it equals three, three in a row. Same thing with columns. Uh, diagonal. So you got to check the diagonal. That's what all this stuff does here is it just goes through the different areas. Actually, let me go ahead and put There we go. Um, and then it comes down here and we do our sum and it says true uh minus three. It just says did somebody win or is it a tie? So, you got to add up all the numbers on there anyway just in case they're all filled up.

And next, we also need to know available positions. Um, these are ones that don't no one's ever used before. This way, when you try something or the computer tries something, uh, it's not going to give it an illegal move. That's what the available positions is doing. Uh, then we want to update our state. And so you have your position going in. We're just sending in the position that you just chose. And you'll see there's a little user interface we put in there. You pick pick the row and column in there. And again, I mean, this is a lot of code. Uh, so really it's kind of a thing you'd want to go through and play with a little bit and just read through it, get a copy of it. Uh, great way to understand how this works. And here is a given reward. Um, so we're going to give a reward. Result equals self-winner. This is one of the hearts of what's going on here. Uh, is we have a result self.winner. So if there's a winner, then we have a result. If the result equals one, here's our feedback. Uh, if it doesn't equal one, then it gets a zero. So it only gets a reward in this particular case if it wins. And that's important to know because different uh systems of reinforced learning do rewarding a lot differently depending on what you're trying to do. This is a very simple example with a 3x3 board. Imagine if you're playing a video game. Uh, certainly you only have so many actions, but your environment is huge; you have a lot going on in the environment, and suddenly a reward system like this is going to be just um is going to have to change a little bit. It's going to have to have different rewards and different setup, and there's all kinds of advanced ways to do that as far as weighing you add weights to it, and so they can add the weights up depending on where the reward comes in. So it might be that you actually get a reward; in this case, you get the reward at the end of the game. And I'm spending just a little bit of time on this because this is an important thing to note, but there's different ways to add up those rewards. It might have like if you take a certain path, um the first reward is going to be weighed a little bit less than the last reward because the last reward is actually winning the game or scoring or whatever it is. So, this reward system gets really complicated on some of the more advanced uh setups.

Um, in this case though, you can see right here that they give a um a 0.1 and a 0.5 reward um just for getting a picking the right value and something that's actually valid instead of picking an invalid value. So rewards again that's like key that's huge. How do you feed the rewards back in? Um, then we have a board reset. That's pretty straightforward. It just goes back and resets the board to the beginning because it's going to try out all these different things while it's learning. It's going to do it by trial and error. So, you have to keep resetting it. And then, of course, there's the play. We want to go ahead and play uh rounds equals 100. Depends on what you want to do on here. Um, you can set this different. You obviously set that to higher level, but this is just going to go through and you'll see in here uh that we have player one and player two. This is this is the computer playing itself. Uh, one of the more powerful ways to learn to play a game or even learn something that isn't a game is to have two of these models that are basically trying to beat each other. And so they always they keep finding explore new things. This one works for this one, so this one tries new things. It beats this. We've seen this in um chess I think was a big one where they had the two players in chess with reinforcement learning uh was one of the ways they train one of the top um computer chess playing algorithms. Uh, so this is just what this is. It's going to choose an action. It's going to try something, and the more it tries stuff um the more we're going to record the hash. We actually have a board hash where you self get the hash set up on here where it stores all the information. And then once you get to a win, one of them wins, it gets the reward. Uh, then we go back and reset and try again. And then kind of the fun part we actually get down here is uh we're going to play with a human. So we'll get a chance to come in here and see what that looks like when you put your own information in. And then it just comes in here and does the same thing it did above. It gives it a reward for its things um or sees if it wins or ties. um looks at available positions, all that kind of fun stuff. And then finally, we want to show the board. Uh, so it's going to print the board out each time really um as an integration is not that exciting. What's exciting uh in here is one looking at this reward system. Whoops. Play one more up. The reward system is really the heart of this. How do you reward the different uh setup? And the other one is when it's playing, it's got to take an action. And so what it chooses for an action is also the heart of reinforcement learning. How do we choose that action? And those are really key to right now where reinforcement learning is um in today's uh technology is uh figuring this out. How do we reward it and how do we guess the next best action?

So we have our uh environment and you can see the environment is we're going to be or the state uh which is kind of like what's going on. We're going to return the state depending on what happens and we want to go ahead and create our agent uh in this case our player. So each one is let me go and grab that. And so we look at a class player. Um this is where a lot of the magic is really going on is what how is this player figuring out how to maneuver around the board? And then the board of course returns a state uh that it can look at and a reward. Uh so we want to take a look at this. We have a name uh self state. This is class player. And when you say class player, we're not talking about a human player. We're talking about um just a uh the computer players. And this is kind of interesting. So remember I told you depending on what you're doing, there's going to be a decay gamma um explore rate. Uh these are what I'm talking about is how do we train it? Um as you try different moves, it gets to the end. The first move is important, but it's not as as important as the last one. And so you could say that u the last one has the heaviest weight. And then as you as you get there, the first one, let's see, the first move gives you a five reward, the second gives you a two reward, and the third one gives you a 10 reward because that's the final ending. You got it; the 10's going to count more than the first step. Uh, and here's our uh, we're going to, you know, get the board information coming in and then choose an action. This was the second part that I was talking about that was so important. Uh, so once you have your training going on, we have to do a little randomness. And you can see right here is our NP random uh, uniform. So it's picking out a random number. Take a random action. This is going to just pick which row and which column it is. Um, and so choosing the action, this one, you can see we're just doing random states, uh, choice, length of positions, action position, and then it skips in there and takes a look at the board, uh, for P and positions. You get, it's actually storing the different boards each time you go through. So, it has a record of what it did so it can properly weigh the values. And this simply just appends a hash state. What's the last date? Pinned it to the uh um to our states on here. Here's our feedback reward. So the reward comes in and it's going to take a look at this and say is it none? Uh, what is the reward? And here is that formula remember I was telling you about up here um that was important because it has decay gamma times the reward. This is where as it goes through each step and this is really important. This is this is kind of the heart of this of what I was talking about earlier. Uh you have step one and this might have a a reward of two. You have step two. I probably should have done ABC. This has a step three. Uh, step four. So on till you get to step n, and this might have a reward of 10. Uh, so reward of 10. We're going to add that, but we're not adding uh let's say this one right here. Uh, let's say this reward here right before 10 was um let's say it's also 10. That just makes the the uh math easy. So we had 10 and 10. Uh we had 10. This is 10 and 10 n whatever it is. But it's time it's 0.9. Uh, so instead of putting a full 10 here, we only do nine. That's uh 0.9 * 10. And so this formula um as far as the decay times the reward minus the cell state value uh it basically adds in it says here's one or here's two. I'm sorry. I should have done this ABC. It would have been easier. Uh, so the first move goes in here and it puts two in here. Uh, then we have our self uh setup on here. You can see how this gets pretty complicated in the math, but this is really the key is how do we train our states and we want the the final state, the win to get the most points. If you win, you get most points. U and the first step gets the least amount of points. So you're really training this almost in reverse. You're retraining you're training it from the last place where you have like it says okay this is now where need to sum up my rewards and I want to sum them up going in reverse and I want to find the answer in reverse kind of an interesting uh uh play on the mind when you're trying to figure this stuff out. And of course, we want to go ahead and reset the board down here. Uh, save the policy, load policy, these are the different things that are going in between the agent and the state to figure out what's going on. Let's go ahead load that up.

And then finally, we want to go ahead and create a human player. And the human player is going to be a little different uh in that uh you choose an action row and column. Here's your action. Uh, if action is if action in positions, meaning positions that are available, uh you return the action. If not, it just keeps asking you until you get an action that actually works. And then we're going to go ahead and append to the hash state which uh we don't need to worry about because it returns the action up here and feed forward. Uh, again this is because it's a human um at the end of the game bat propagate and update state values. This part isn't being done because it's not programming uh the model. Uh, the model is getting its own rewards. So we've gone ahead and loaded this in here. Uh, so here's all our pieces. And the first thing we want to do is set up uh P1 player one, uh, P2, player two, and then we're going to send our players to our state. So now it has P1, P2, and it's going to play, and it's going to play 50,000 rounds. Now, we can probably do a lot less than this, and it's not going to get the full results. In fact, you know what? Uh, let's go ahead and just do five. Uh, just to play with it because I want to show you something here. Oops. Somewhere in there I forgot to load something. There we go. I must have forgot to run this run. Oops, forgot a reference there for the board rows and columns 3x3. Um, there is actually in the state it references that. We just tack it on on the end. It was supposed to be at the beginning. Uh, so now I've only set this up with um, see where are we going here? I've only set this up to train five times. And the reason I did that is we're going to uh, come

In and actually play it. And then I'm going to change that, and we can see how it differs on there. There we go. And it didn't make it through a run. And we're going to go ahead and save the policy. Um, so now we have our player one and our player two policy. Uh, the way we set it up, it has two separate policies loaded up in there.

And then we're going to come in here, and we're going to do uh player one is going to be the computer experience rate zero. Load policy one. Human player human. And we're going to go ahead and play this. Now remember, I only went through it um uh just one round of training. In fact, minimal training. And so it puts an X there. And I'm going to go ahead and do row zero, column one. You can see this is very uh basic on here. And so I put in my zero. And then I'm going to go zero, block it, zero, zero. And you can see right here, it let me win. Uh, just like that, I was able to win. Zero two. And woo, human wins. So I only trained it five times. We're going to run this again. And this time, uh, instead of five, let's do 5,000 or 50,000. I think that's what the guys in the back had. And this takes a while to train it. This is where reinforcement learning really falls apart. Look how simple this game is. We're talking about a 3x3 set of columns. And so for me to train it on this um I could do a Q table which would take which would go much quicker. Um, you could build a quick Q table with almost all the different options on there, and uh you would probably get a the same result much quicker. We're just using this as an example.

So when we look at reinforcement learning, you need to be very careful what you apply it to. It sounds like a good deal until you do like a large neural network where you're doing um you set the neural network to a learning increment of one. So every time it goes through it learns, and then you do your action. So you pick from the learning uh setup, and you actually try actions on the learning setup until you get the what you think is going to be the best action. So you actually feed what you think is right back through the neural network. There's a whole layer there which is really fun to play with, and then it has an output. Well, think of all those processes. I mean, that is just a huge amount of work it's going to do. Uh, let's go ahead and skip ahead here. Give it a moment. It's going to take a a minute or two to go ahead and run now to train it. Uh, we went ahead and let it run, and it took a while. This this took um I got a pretty powerful processor, and it took about five minutes plus to run it, and we'll go ahead and uh run our player setup on here. Oops, it brought in the last Whoops, it brought in the last round. So, give me just a moment to reddo the policy save. There we go. I forgot to save the policy back in there and then go ahead and run our player again.

So, we've saved the policy, and then we want to go ahead and load the policy for P1 as a computer. And we can see the computer's gone in the bottom right corner. I'm going to go ahead and go uh one one which is the center, and it's gone right up the top. And if you have ever played tic-tac-toe, you know the computer has me. Uh, but we'll go ahead and play it out. Row zero, column two. There it is. And then it's gone here. And so I'm going to go ahead and go row 0 one two. No, 01. There we go. And column zero. That's where I wanted. Oh, and it says I Okay, you your action. There we go. Boom. Uh, so you can see here we've got a Didn't catch the win on this. It said tie. Um, kind of funny that it didn't catch the win on there. But if we play this a bunch of times, you'll find it's going to win more and more. The more we train it, the more the reinforcement happens. This lengthy training process uh is really the stopper on reinforcement learning. As this changes, reinforcement learning will be one of the more powerful uh packages evolving over the next decade or two. In fact, I would even go as far as to say it is the most important uh machine learning tool and artificial intelligence tool out there as it learns not only a simple tic-tac-toe board, but we start learning environments. And the environment would be like in language. If you're translating a language or something from one language to the other, so much of it is lost if you don't know the context, it's in what's the environments it's in. And so being able to attach environment and context and all those things together is going to require reinforcement learning to do. So again, if you want to get a copy of the tic-tac-toe board, it's kind of fun to play with. Uh run it, you can test it out, you can do u you know, test it for different uh uh values. You can switch from P1 computer uh where we loaded the policy one to load the policy two and just see how it varies. There's all kinds of things you can do on there.

Supervised learning uses labeled data to train machine learning models. Label data means that the output is already known to you. The model just needs to map the inputs to the outputs. An example of supervised learning can be to train a machine that identifies the image of an animal. Below you can see we have a trained model that identifies the picture of a cat.

Unsupervised learning uses unlabelled data to train machines. Unlabeled data means there is no fixed output variable. The model learns from the data, discovers patterns and features in the data and returns the output. Here is an example of an unsupervised learning technique that uses the images of vehicles to classify if it's a bus or a truck. So the model learns by identifying the parts of a vehicle such as the length and width of the vehicle, the front and rear end covers, roof hoods, the types of wheels used, etc. Based on these features, the model classifies if the vehicle is a bus or a truck.

Reinforcement learning trains a machine to take suitable actions and maximize reward in a particular situation. It uses an agent and an environment to produce actions and rewards. The agent has a start and an end state, but there might be different parts for reaching the end state like a maze. In this learning technique, there is no predefined target variable. An example of reinforcement learning is to train a machine that can identify the shape of an object given a list of different objects such as square, triangle, rectangle or a circle. In the example shown, the model tries to predict the shape of the object which is a square.

Here now let's look at the different machine learning algorithms that come under these learning techniques. Some of the commonly used supervised learning algorithms are linear regression, logistic regression, support vector machines, K nearest neighbors, decision tree, random forest and naive Bayes. Examples of unsupervised learning algorithms are K means clustering, hierarchical clustering, DB scan, principal component analysis and others. Choosing the right algorithm depends on the type of problem you are trying to solve. Some of the important reinforcement learning algorithms are Q-learning, Monte Carlo, SARSA and deep Q network.

Now let's look at the approach in which these machine learning techniques work. So supervised learning takes labeled inputs and maps it to known outputs which means you already know the target variable. Unsupervised learning finds patterns and understands the trends in the data to discover the output. So the model tries to label the data based on the features of the input data. While reinforcement learning follows trial and error method to get the desired solution. After accomplishing a task, the agent receives an award. An example could be to train a dog to catch the ball. If the dog learns to catch a ball, you give it a reward such as a biscuit.

Now let's discuss the training process for each of these learning methods. So supervised learning methods need external supervision to train machine learning models and hence the name supervised. They need guidance and additional information to return the result. Unsupervised learning techniques do not need any supervision to train models. They learn on their own and predict the output. Similarly, reinforcement learning methods do not need any supervision to train machine learning models.

And with that let's focus on the types of problems that can be solved using these three types of machine learning techniques. So supervised learning is generally used for classification and regression problems. We'll see the examples in the next slide. And unsupervised learning is used for clustering and association problems. While reinforcement learning is reward based. So for every task or for every step completed there will be a reward received by the agent. And if the task is not achieved correctly, there will be some penalty used.

Now let's look at a few applications of supervised, unsupervised and reinforcement learning. As we saw earlier, supervised learning are used to solve classification and regression problems. For example, you can predict the weather for a particular day based on humidity, precipitation, wind speed, and pressure values. You can use supervised learning algorithms to forecast sales for the next month or the next quarter for different products. Similarly, you can use it for stock price analysis or identifying if a cancer cell is malignant or benign.

Now talking about the applications of unsupervised learning, we have customer segmentation. So based on customer behavior, likes, dislikes and interests, you can segment and cluster similar customers into a group. Another example where unsupervised learning algorithms are used is customer churn analysis.

Now let's see what applications we have in reinforcement learning. So reinforcement learning algorithms are widely used in the gaming industries to build games. It is also used to train robots to perform human tasks. If you are an aspiring data scientist who's looking out for online training and certification in data science from the best universities and industry experts, then search no more. Simply learns post-graduate program in data science from Caltech University in collaboration with IBM should be the right choice. For more details on this program, please use the link in the description box below.

Often professionals want to know if there is a relationship between two or more variables. For instance, is there a relationship between the grade on the third French exam a student takes and the grade on the final exam? If yes, then how is it related and how strongly? Regression can be used here to arrive at a conclusion. This is an example of bivariate data that is two variables. However, statisticians are mostly interested in multivariate data. Regression analysis is used to predict the value of one variable, the dependent variable, on the basis of other variables, the independent variables. In the simplest form of regression, linear regression, you work with one independent variable. The formula for simple linear regression is shown on the screen. In the next screen, we'll look at a few examples of regression analysis. Regression analysis is used in several situations such as those described on the screen. In example one, using the data given on the screen, you have to analyze the relation between the size of a house and its selling price for a realer. In example two, you need to predict the exam scores of students who study for 7.2 hours with the help of the data shown on the slide. A couple more examples are given on the screen. In example three, based on the expected number of customers and the previous day's data given, you need to predict the number of burgers that will be sold by a KFC outlet. In example four, you have to calculate the life expectancy for a group of people with the average length of schooling based on the data given.

Let's look at the two main types of regression analysis. Simple linear regression and multiple linear regression. Both of these statistical methods use a linear equation to model the relationship between two or more variables. Simple linear regression considers one quantitative and independent variable X to predict the other quantitative but dependent variable Y. Multiple linear regression considers more than one quantitative and qualitative variable to predict a quantitative and dependent variable Y. We'll look at the two types of analyses in more detail in the slides that follow.

In simple linear regression, the predictions of the explained variable Y when plotted as a function of the explanatory variable X form a straight line. The best fitting line is called the regression line. The output of this model is a function to predict the dependent variable on the basis of the values of the independent variable. The dependent variable is continuous and the independent variable can be continuous or discrete. Let's look at the different kinds of linear and nonlinear analyses. List of linear techniques are simple method of least squares, coefficient of multiple determination, standard error of the estimate, dummy variable and interaction. Similarly, there are many nonlinear techniques available such as polynomial, logarithmic, square root, reciprocal and exponential.

To understand this model, we'll first look at a few assumptions. The simple linear regression model depicts the relationship between one dependent and two or more independent variables. The assumptions which justify the use of this model are as follows: Linear and additive relationship between the dependent and independent variables; Multivariate normality; Little or no colinearity in the data; Little or no autocorrelation in the data; Homoscedasticity, that is variance of errors same across all values of X. The equation for this model is shown on the screen. A more descriptive graphical representation of simple linear regression is given on the screen. Beta KN represents the slope. A slope with two variables implies that one unit changes in X result in a two-unit change in Y. Beta 1 represents the estimated change in the average value of Y as a result of one unit change in X. Epsilon represents the estimated average value of Y when the value of X is zero.

This demo will show the steps to do simple linear regression in R. In this demo, you'll learn how to do simple linear regression. Let's use X and Y vectors that we have created in the previous demo. We also ensured there exist a relationship between X and Y visually by plotting a graph. To build a simple linear regression model, let's use the LM function. To see how the linear model fits into X and Y, let's plot the linear line by the AB line function. Let's use the predict function to test or predict the linear model. We can pass a known variable to predict the unknown variables.

Let's look at an example of a common use for linear regression: Profit estimation of a company. If I was going to invest in a company, I would like to know how much money I could expect to make. So, we'll take a look at a venture capitalist firm and try to understand which companies they should invest in. So we'll take the idea that we need to decide the companies to invest in. We need to predict the profit the company makes, and we're going to do it based on the company's expenses and even just a specific expense. In this case we have our company, we have the different expenses. So we have our R&D which is your research and development. We have our marketing. Uh we might have the location. We might have what kind of administration it's going through. Based on all this different information we would like to calculate the profit. Now, in actuality, there's usually about 23 to 27 different markers that they look at if they're a heavy-duty investor. We're only going to take a look at one basic one. We're going to come in and for simplicity, let's consider a single variable, R&D, and find out which companies to invest in based on that. So, we take our R&D and we're plotting the profit based on the R&D expenditure, how much money they put into the research and development. And then we look at the profit that goes with that. We can predict a line to estimate the profit. So we can draw a line right through the data. And when you look at that, you can see how much they invest in the R&D is a good marker as to how much profit they're going to have. We can also note that companies spending more on R&D make good profit. So let's invest in the ones that spend a higher rate in their R&D.

What's in it for you? First, we'll have an introduction to machine learning followed by machine learning algorithms. These will be specific to linear regression and where it fits into the larger model. Then we'll take a look at applications of linear regression, understanding linear regression, and multiple linear regression. Finally, we'll roll up our sleeves and do a little programming in use case profit estimation of companies. Let's go ahead and jump in. Let's start with our introduction to machine learning along with some machine learning algorithms and where that fits in with linear regression.

Let's look at another example of machine learning. Based on the amount of rainfall, how much would be the crop yield? So we here we have our crops, we have our rainfall, and we want to know how much we're going to get from our crops this year. So we're going to introduce two variables, independent and dependent. The independent variable is a variable whose value does not change by the effect of other variables and is used to manipulate the dependent variable. It is often denoted as X. In our example, rainfall is the independent variable. This is a wonderful example because you can easily see that we can't control the rain but the rain does control the crop. So we talk about the independent variable controlling the dependent variable. Let's define dependent variable as a variable whose value change when there is any manipulation in the values of the independent variables. It is often denoted as y. And you can see here our crop yield is dependent variable and it is dependent on the amount of rainfall received.

Now that we've taken a look at a real-life example, let's go a little bit into the theory and some definitions on machine learning and see how that fits together with linear regression. Numerical and categorical values. Let's take our data coming in, and this is kind of random data from any kind of project. We want to divide it up into numerical and categorical. So numerical is numbers: age, salary, height, where categorical would be a description: the color, a dog's breed, gender. Categorical is limited to very specific items where numerical is a range of information.

Now that you've seen the difference between numerical and categorical data, let's take a look at some different machine learning definitions. When we look at our different machine learning algorithms, we can divide them into three areas: supervised, unsupervised, reinforcement. We're only going to look at supervised today. Unsupervised means we don't have the answers and we're just grouping things. Reinforcement is where we give positive and negative feedback to our algorithm to program it, and it doesn't have the information till after the fact. But today we're just looking at supervised because that's where linear regression fits in. In supervised data, we have our data already there and our answers for a group, and then we use that to program our model and come up with an answer. The two most common uses for that is through the regression and classification. Now we're doing linear regression. So we're just going to focus on the regression side. And in the regression we have simple linear regression, we have multiple linear regression, and we have polynomial linear regression. Now on these three simple linear regression is the examples we've looked at so far where we have a lot of data and we draw a straight line through it. Multiple linear regression means we have multiple variables. Remember where we had the rainfall and the crops? We might add additional variables in there like how much food do we give our crops? When do we harvest them? Those would be additional information added to our model, and that's why it'd be multiple linear regression. And finally we have polynomial linear regression. That is instead of drawing a line we can draw a curved line through it.

Now that you see where regression model fits into the machine learning algorithms and we're specifically looking at linear regression. Let's go ahead and take a look at applications for linear regression. Let's look at a few applications of linear regression: Economic growth used to determine the economic growth of a country or a state in the coming quarter, can also be used to predict the GDP of a country. Product price can be used to predict what would be the price of a product in the future. We can guess whether it's going to go up or down or should I buy today. Housing sales to estimate the number of houses a builder would sell and what price in the coming months. Score predictions. Cricket fever to predict the number of runs a player would score in the coming matches based on the previous performance. I'm sure you can figure out other applications you could use linear regression for. So, let's jump in and let's understand linear regression and dig into the theory.

Understanding linear regression. Linear regression is the statistical model used to predict the relationship between independent and dependent variables by examining

Two factors. The first important one is which variables in particular are significant predictors of the outcome variable. And the second one that we need to look at closely is how significant is the regression line to make predictions with the highest possible accuracy. If it's inaccurate, we can't use it. So it's very important we find out the most accurate line we can get.

Since linear regression is based on drawing a line through data, we're going to jump back and take a look at some uklitian geometry. The simplest form of a simple linear regression equation with one dependent and one independent variable is represented by y = m * x + c. And if you look at our model here, we plotted two points on here. Uh, x1 and y1, x2 and y2. y being the dependent variable, remember that from before. And x being the independent variable. So y depends on whatever x is. M in this case is the slope of the line where M equals the difference in the Y2 - Y1 and X2 - X1. And finally, we have C which is the coefficient of the line or where it happens to cross the zero axis.

Let's go back and look at an example we used earlier of linear regression. We're going to go back to plotting the amount of crop yield based on the amount of rainfall. And here we have our rainfall. Remember, we cannot change rainfall. And we have our crop yield, which is dependent on the rainfall. So, we have our independent and our dependent variables. We're going to take this and draw a line through it as best we can through the middle of the data. And then we look at that. We put the red point on the y axis is the amount of crop yield you can expect for the amount of rainfall represented by the green dot. So, if we have an idea what the rainfall is for this year and what's going on, then we can guess how good our crops are going to be. And we've created a nice line right through the middle to give us a nice mathematical formula. Let's take a look and see what the math looks like behind this.

Let's look at the intuition behind the regression line. Now, before we dive into the math and the formulas that go behind this and what's going on behind the scenes, I want you to note that when we get into the case study and we actually apply some Python script that this math that you're going to see here is already done automatically for you. You don't have to have it memorized. It is, however, good to have an idea what's going on so if people reference the different terms, you'll know what they're talking about.

Let's consider a sample data set with five rows and find out how to draw the regression line. We're only going to do five rows because if we did like the rainfall with hundreds of points of data, that would be very hard to see what's going on with the mathematics. So, we'll go ahead and create our own two sets of data. And we have our independent variable x and our dependent variable y. And when x was 1, we got y = 2. When x was uh 2, y was 4. And so on and so on. If we go ahead and plot this data on a graph, we can see how it forms a nice line through the middle. You can see where it's kind of grouped going upwards to the right. The next thing we want to know is what the means is of each of the data coming in, the x and the y. The means doesn't mean anything other than the average. So we add up all the numbers and divide by the total. So 1 + 2 + 3 + 4 + 5 over 5 equals three. And the same for y, we get four. If we go ahead and plot the means on the graph, we'll see we get 3, 4, which draws a nice line down the middle, a good estimate.

Here we're going to dig deeper into the math behind the regression line. Now, remember before I said you don't have to have all these formulas memorized or fully understand them, even though we're going to go into a little more detail of how it works. And if you're not a math wiz and you don't know if you've never seen the sigma character before, which looks a little bit like an e that's opened up, that just means summation. That's all that is. So, when you see the sigma character, it just means we're adding everything in that row. And for computers, this is great because as a programmer, you can easily iterate through each of the XY points and create all the information you need. So in the top half, you can see where we've broken that down into pieces. And as it goes through the first two points, it computes the squared value of X, the squared value of Y, and X * Y. And then it takes all of X and adds them up. All of Y adds them up. All of X squared, adds them up, and so on and so on. And you can see we have the sum of equal to 15. The sum is equal to 20. All the way up to x * y where the sum equals 66. This all comes from our formula for calculating a straight line where y equals the slope * x plus the coefficient c. So we go down below and we're going to compute more like the averages of these. And we're going to explain exactly what that is in just a minute and where that information comes from. It's called the square means error, but we'll go into that in detail in a few minutes. All you need to do is look at the formula and see how we've gone about computing it line by line instead of trying to have a huge set of numbers pushed into it. And down here you'll see where the slope m equals and then the top part if you read through the brackets you have the number of data points times the sum of x * y which we computed one line at a time there. And that's just the 66. And take all that and you subtract it from the sum of x times the sum of y and those have both been computed. So you have 15 * 20 and on the bottom we have the number of lines times the sum of x^2 easily computed as 86 for the sum minus I'll take all that and subtract the sum of x^2 and we end up as we come across with our formula you can plug in all those numbers which is very easy to do on the computer you don't have to do the math on a piece of paper or calculator and you'll get a slope of 6 and you'll get your C coefficient if you continue to follow through that formula you'll see it comes out as equal to 2.2.

Continuing deeper into what's going behind the scenes, let's find out the predicted values of y for corresponding values of x using the linear equation where m=6 and c = 2.2. We're going to take these values and we're going to go ahead and plot them. We're going to predict them. So y = 6 * x = 1 + 2.2 = 2.8 so on and so on. And here the blue points represent the actual y values and the brown points represent the predicted y values based on the model we created. The distance between the actual and predicted values is known as residuals or errors. The best fit line should have the least sum of squares of these errors also known as equare. If we put these into a nice chart where you can see X and you can see Y what the actual values were and you can see Y predicted you can easily see where we take Y minus Y predicted and we get an answer what is the difference between those two and if we square that Y minus Y prediction squared we can then sum those squared values that's where we get the 64 plus the .36 + 1 all the way down until we have a summation equals 2.4. So the sum of squared errors for this regression line is 2.4. We check this error for each line and conclude the best fit line having the least e value. Get a nice graphical representation. We can see here where we keep moving this line through the data points to make sure the best fit line has the least squared distance between the data points and the regression line. Now we only looked at the most commonly used formula for minimizing the distance. There are lots of ways to minimize the distance between the line and the data points like sum of squared errors, sum of absolute errors, root mean square error, etc. What you want to take away from this is whatever formula is being used, you can easily using a computer programming and iterating through the data calculate the different parts of it. That way, these complicated formulas you see with the different summations and absolute values are easily computed one piece at a time.

Up until this point, we've only been looking at two values, X and Y. Well, in the real world, it's very rare that you only have two values when you're figuring out a solution. So, let's move on to the next topic, multiple linear regression.

Let's take a brief look at what happens when you have multiple inputs. So, in multiple linear regression, we have uh, well, we'll start with the simple linear regression where we had y = m + x + c and we're trying to find the value of y. Now with multiple linear regression we have multiple variables coming in. So instead of having just x we have x1 x2 x3 and instead of having just one slope each variable has its own slope attached to it. As you can see here we have m1 m2 m3 and we still just have the single coefficient. So when you're dealing with multiple linear regression you basically take your single linear regression and you spread it out. So you have y = m1 * x1 + m2 * x2 so on all the way to m to the nth x to the nth and then you add your coefficient on there.

Implementation of linear regression. Now we get into my favorite part. Let's understand how multiple linear regression works by implementing it in Python. If you remember before we were looking at a company and just based on its R&D trying to figure out its profit. We're going to start looking at the expenditure of the company. We're going to go back to that. We're going to predict its profit, but instead of predicting it just on the R&D, we're going to look at other factors like administration costs, marketing costs, and so on. And from there, we're going to see if we can figure out what the profit of that company's going to be.

To start our coding, we're going to begin by importing some basic libraries. And we're going to be looking through the data before we do any kind of linear regression. We're going to take a look at the data to see what we're playing with. Then we'll go ahead and format the data to the format we need to be able to run it in the linear regression model and then from there we'll go ahead and solve it and just see how valid our solution is. So let's start with importing the basic libraries.

Now I'm going to be doing this in Anaconda Jupyter notebook, a very popular IDE. I enjoy it cuz it's such a visual to look at and so easy to use. Um, just any IDE for Python will work just fine for this. So break out your favorite Python IDE. So, here we are in our Jupyter notebook. Let me go ahead and paste our first piece of code in there. And let's walk through what libraries we're importing. First, we're going to import numpy as np. And then I want you to skip one line and look at import pandas as pd. These are very common tools that you need with most of your linear regression. The numpy, which stands for number python, is usually denoted as np and you have to almost have that for your sklearn toolbox. So, you always import that right off the beginning. pandas. Although you don't have to have it for your sklearn libraries, it does such a wonderful job of importing data, setting it up into a data frame so we can manipulate it rather easily and it has a lot of tools also in addition to that. So we usually like to use the pandas when we can and I'll show you what that looks like. The other three lines are for us to get a visual of this data and take a look at it. So we're going to import mapplot library.pyplot as plt and then seaborn as sns. Seaborn works with the matplotlib library. So you have to always import matplotlib and then seaborn sits on top of it. And we'll take a look at what that looks like. You could use any of your own plotting libraries you want. There's all kinds of ways to look at the data. These are just very common ones. And the seaborn is so easy to use. It just looks beautiful. It's a nice representation that you can actually take and show somebody. And the final line is the %matplotlib inline. That is only because I'm doing an inline IDE. My interface in the Anaconda Jupyter notebook requires I put that in there or you're not going to see the graph when it comes up. Let's go ahead and run this. It's not going to be that interesting because we're just setting up variables. In fact, it's not going to do anything that we can see, but it is importing these different libraries and setup.

The next step is load the data set and extract independent and dependent variables. Now, here in the slide, you'll see companies = pd.read_csv. And it has a long line there with the file at the end. 1000_companies.csv. You're going to have to change this to fit whatever setup you have. And the file itself, you can request. Just go down to the commentary below this video and put a note in there and SimplyLearn will try to get in contact with you and supply you with that file so you can try this coding yourself. So, we're going to add this code in here. And we're going to see that I have companies = pd.read_csv. And I've changed this path to match my computer. C:/simplylearn/1000_companies.csv. And then below there, we're going to set the x = to companies.iloc. And because this is companies is a pd data set, I can use this nice notation that says take every row, that's what the colon, the first colon is, comma, except for the last column. That's what the second part is where we have a colon -1 and we want the values set into there. So X is no longer a data set, a Pandas data set, but we can easily extract the data from our pandas data set with this notation. And then Y we're going to set equal to the last row. Well, the question is going to be what are we actually looking at? So let's go ahead and take a look at that. And we're going to look at the companies.head which lists the first five rows of data. And I'll open up the file in just a second so you can see where that's coming from. But let's look at the data in here as far as the way the pandas sees it. When I hit run, you'll see it breaks it out into a nice setup. This is what pandas, one of the things pandas is really good about is it looks just like an Excel spreadsheet. You have your rows and remember when we're programming, we always start with zero. We don't start with one. So it shows the first five rows, 0, 1, 2, 3, 4. And then it shows your different columns. R&D spend, administration, marketing spend, state, profit. It even notes that the top are column names. It was never told that, but Pandas is able to recognize a lot of things that they're not the same as the data rows. Why don't we go ahead and open this file up in a CSV so you can actually see the raw data. So here I've opened it up as a text editor. And you can see at the top we have R&D spend, administration, marketing spend, state, profit, carriage return. I don't know about you, but I'd go crazy trying to read files like this. That's why we use the pandas. You could also open this up in an Excel and it would separate it since it is a comma-separated variable file. But we don't want to look at this one. We want to look at something we can read rather easily. So let's flip back and take a look at that top part, the first five rows.

Now, as nice as this format is where I can see the data, to me it doesn't mean a whole lot. Maybe you're an expert in business and investments and you understand what $165,349.20 compared to the administration cost of $136,897.80 so on so on helps to create the profit of $192,261.83. That makes no sense to me whatsoever. No pun intended. So let's flip back here and take a look at our next set of code where we're going to graph it so we can get a better understanding of our data and what it means. So at this point we're going to use a single line of code to get a lot of information so we can see where we're going with this. Let's go ahead and paste that into our uh notebook and see what we got going. And so we have the visualization and again we're using sns which is pandas. As you can see we imported the matplotlib.pyplot as plt which then the seaborn uses and we imported the seaborn as sns and then that final line of code helps us show this in our inline coding without this it wouldn't display and you can display it to a file and other means and that's the %matplotlib inline with the amberite at the beginning. So here we come down to the single line of code seaborn is great because it actually recognizes the panda data frame so I can just take the companies.corr for coordinates and I can put that right into the seaborn. And when we run this, we get this beautiful plot. And let's just take a look at what this plot means. If you look at this plot on mine, the colors are probably a little bit more purplish and blue than the original one. Uh, we have the columns and the rows. We have R and D spending, we have administration, we have marketing spending, and profit. And if you cross index any two of these, since we're interested in profit, if you cross index profit with profit, it's going to show up, if you look at the scale on the right, way up in the dark. Why? Because those are the same data. They have an exact correspondence. So R&D spending is going to be the same as R&D spending and the same thing with administration costs. So right down the middle, you get this dark row or dark um diagonal row that shows that this is the highest corresponding data that's exactly the same. And as it becomes lighter, there's less connections between the data. So we can see with profit, obviously profit is the same as profit. And next, it has a very high correlation with R&D spending, which we looked at earlier, and it has a slightly less connection to marketing spending, and even less to how much money we put into the administration. So, now that we have a nice look at the data, let's go ahead and dig in and create some actual useful linear regression models so that we can predict values and have a better profit.

Now that we've taken a look at the visualization of this data, we're going to move on to the next step. Instead of just having a pretty picture, we need to generate some hard data, some hard values. So let's see what that looks like. We're going to set up our linear regression model in two steps. The first one is we need to prepare some of our data so it fits correctly. And let's go ahead and paste this code into our Jupyter notebook. And what we're bringing in is we're going to bring in the sklearn.preprocessing where we're going to import the LabelEncoder and the OneHotEncoder. To use the LabelEncoder, we're going to create a variable called label_encoder and set it equal to LabelEncoder(). This creates a class that we can reuse for transferring the labels back and forth. Now, about now, you should ask, what labels are we talking about? Let's go take a look at the data we processed before and see what I'm talking about here. If you remember when we did the companies.head and we printed the top five rows of data, we have our columns going across. We have column zero which is R&D spending, column one which is administration, column two which is marketing spending and column three is state and you'll see under state we have New York, California, Florida. Now to do a linear regression model, it doesn't know how to process New York. It knows how to process a number. So the first thing we're going to do is we're going to change that New York, California, and Florida.

and we're going to change those to numbers. That's what this line of code does here: X equals and then it has the colon, comma, 3 in brackets. The first part, the colon, comma, means that we're going to look at all the different rows. So, we're going to keep them all together, but the only row we're going to edit is the third row. And in there, we're going to take the label coder and we're going to fit and transform the X, also the third row. So, we're going to take that third row, we're going to set it equal to a transformation. And that transformation basically tells it that instead of having a uh New York, it has a zero or a one or a two.

And then finally we need to do a one hot encoder which equals one hot in order categorical features equals three. And then we take the X and we go ahead and do that equal to one hot encoder fit transform X to array. This final transformation preps our data for us. So it's completely set the way we need it as just a row of numbers, even though it's not in here. Let's go ahead and print x and just take a look what this data is doing. You'll see I have an array of arrays and then each array is a row of numbers. And if I go ahead and just do row zero, you'll see I have a nice organized row of numbers that the computer now understands. We'll go ahead and take this out there because it doesn't mean a whole lot to us. It's just a row of numbers.

Next on setting up our data, we have avoiding dummy variable trap. This is very important. Why? Because the computers automatically transformed our header into the setup and it's automatically transformed all these different variables. So when we did the encoder, the encoder created two columns and what we need to do is just have the one because it has both the variable and the name. That's what this piece of code does here. Let's go ahead and paste this in here. And we have x= x colon, one colon. All this is doing is removing that one extra column we put in there when we did our one hot encoder and our label encoding. Let's go ahead and run that.

And now we get to create our linear regression model. And let's see what that looks like here. And we're going to do that in two steps. The first step is going to be in splitting the data. Now whenever we create a uh predictive model of data, we always want to split it up. So we have a training set and we have a testing set. That's very important. Otherwise, we'd be very unethical without testing it to see how good our fit is. And then we'll go ahead and create our multiple linear regression model and train it and set it up. Let's go ahead and paste this next piece of code in here. And I'll go ahead and shrink it down a size or two so it all fits on one line. So from the sklearn module selection, we're going to import train test split. And you'll see that we've created four completely different variables: We have capital X train, capital X test, smallerase Y train, smallerase Y test. That is the standard way that they usually reference these when we're doing different uh models. Usually see that a capital X and you see the train and the test and the lowercase Y. What this is is X is our data going in. That's our R&D spin, our administration, our marketing. And then Y, which we're training, is the answer. That's the profit because we want to know the profit of an unknown entity. So that's what we're going to shoot for in this tutorial.

The next part, train test split. We take X and we take Y. We've already created those. X has the columns with the data in it and Y has a column with profit in it. And then we're going to set the test size equals.2. That basically means 20%. So 20% of the rows are going to be tested. We're going to just put them off to the side. So since we're using a thousand lines of data, that means that 200 of those lines we're going to hold off to the side to test for later. And then the random state equals zero, we're going to randomize which ones it picks to hold off to the side. We'll go ahead and run this. It's not overly exciting because it's setting up our variables. But the next step is the next step we actually create our linear regression model.

Now that we got to the linear regression model, we get that next piece of the puzzle. Let's go ahead and put that code in there and walk through it. So here we go. We're going to paste it in there and let's go ahead and uh since this is a shorter line of code, let's zoom up there so we can get a good look. And we have from the sklearn.linear_model, we're going to import linear regression. Now, I don't know if you recall from earlier when we were doing all the math. Let's go ahead and flip back there and take a look at that. Do you remember this where we had this long formula on the bottom and we were doing all this summization and then we also looked at setting it up with the different lines and then we also looked all the way down to multiple linear regression where we're adding all those formulas together? All of that is wrapped up in this one section. So what's going on here is I'm going to create a variable called regressor and the regressor equals the linear regression. That's a linear regression model that has all that math built in. So we don't have to have it all memorized or have to compute it individually. And then we do the regressor.fit. In this case, we do X train and Y train because we're using the training data. X being the data in and Y being profit what we're looking at. And this does all that math for us. So within one click and one line, we've created the whole linear regression model and we fit the data to the linear regression model. And you can see that when I run the regressor, it gives an output linear regression. It says copy x equals true fit intercept equals true in jobs equal 1 normalize equals false. It's just giving you some general information on what's going on with that regressor model.

Now that we've created our linear regression model, let's go ahead and use it. And if you remember, we kept a bunch of data aside. So we're going to do a y predict variable and we're going to put in the x test. And let's see what that looks like. Scroll up a little bit. Paste that in here. Predicting the test set results. So here we have y predict equals regressor.predict x test going in and this gives us y predict. Now because I'm in jupiter in line I can just put the variable up there and when I hit the run button it'll print that array out. I could have just as easily done print y predict. So if you're in a different IDE that's not an inline setup like the Jupyter notebook, you can do it this way: Print y predict. And you'll see that for the 200 different test variables we kept off to the side, it's going to produce 200 answers. This is what it says the profit are for those 200 predictions. But let's don't stop there. Let's keep going and take a couple look. We're going to take just a short detail here and calculating the coefficients and the intercepts. This gives us a quick flash at what's going on behind the line.

We're going to take a short detour here and we're going to be calculating the coefficient and intercepts. So you can see what those look like. What's really nice about our regressor we created is it already has the coefficients for us. We can simply just print regressor.coefficient_. When I run this, you'll see our coefficients here. And if we can do the regressor coefficient, we can also do the regressor intercept. And let's run that and take a look at that. This all came from the multiple regression model. And we'll flip over so you can remember where this is going into, where it's coming from. You can see the formula down here where y = m1 * x1 + m2 * x2 and so on and so on plus c the coefficient. So these variables fit right into this formula: y = slope 1 * column 1 variable plus slope 2 * column 2 variable all the way to the m into the n and x to the n + c the coefficient. Or in this case you have - 8.89 8 9 to the power of two etc etc times the first column and the second column and the third column and then our intercept is the minus one3009 point it gets kind of complicated when you look at it. This is why we don't do this by hand anymore. This is why we have the computer to make these calculations easy to understand and calculate.

Now I told you that was a short detour and we're coming towards the end of our script. As you remember from the beginning, I said if we're going to divide this information, we have to make sure it's a valid model that this model works and understand how good it works. So calculating the R squar value, that's what we're going to use to predict how good our prediction is. And let's take a look at what that looks like in code. And so we're going to use this from sklearn.metrics. We're going to import R2 score. That's the R squared value. We're looking at the error. So in the R2 score, we take our Y test versus our Y predict. Y test is the actual values we're testing. That was the one that was given to us that we know are true. The Y predict of those 200 values is what we think it was true. And when we go ahead and run this, we see we get a 9352. That's the R2 score. Now, it's not exactly a straight percentage, so it's not saying it's 93% correct, but you do want that in the upper 90s 0 and higher shows that this is a very valid prediction based on the R2 score. And if R squared value of N1 or 92 as we got on our model remember it does have a random generation involved. This proves the model is a good model which means success. Yay. We successfully trained our model with certain predictors and estimated the profit of the companies using linear regression.

So now that we have a successful linear regression model, let's take a look at what we went over today and take a look at our key takeaways. First, if you are an aspiring data scientist who's looking out for online training and certification in data science from the best universities and industry experts, then search no more. Simply learns post-graduate program in data science from Caltech University in collaboration with IBM should be the right choice. For more details on this program, please use the link in the description box below.

What is logistic regression? Let's say we have to build a predictive model or a machine learning model to predict whether the passengers of the Titanic ship have survived or not the shipwreck. So how do we do that? So we use logistic regression to build a model for this. How do we use logistic regression? So we have the information about the passengers, their ID, whether they have survived or not, their class and name and so on and so forth. And we use this information where we already know whether the person has survived or not. That is the labeled information. And we help the system to train based on this information based on this labeled data. This is known as label data. And during the process of building the model, we probably will remove some of the non-essential parameters or attributes here. We only take those attributes which are really required to make these predictions. And once we train the model, we run new data through it whereby the model will predict whether the passenger has survived or not.

So let's see what we will learn in this video. We will talk about what is supervised learning and we will go into details about classification which is one of the techniques for supervised learning and then we will further focus on logistic regression is which is one of the algorithms for performing classification especially binary classification. Then we will compare linear and logistic regression and what are some of the logistic regression applications and finally we will end with a use case or a demo of actual Python code for doing logistic regression in Jupyter notebook. All right. So let's start with what is supervised learning. Supervised learning is one of the two main types of machine learning methods. Here we use what is known as labeled data to help the system learn. This is very similar to how we human beings learn. So let's say you want to teach a child to recognize an apple. How do we do that? We never tell the child okay this is an apple has a certain diameter on the top certain diameter at the bottom and this has a certain RGB color. No, we just show an apple to the child and tell the child this is apple and then next time when we show an apple, child immediately recognizes yes this is an apple. Supervised learning works very similar on the similar lines.

So where does logistic regression fit into the overall machine learning process? Machine learning is divided into two types mainly two types. There is a third one called reinforcement learning but we will not talk about that right now. So one is supervised learning and the other is unsupervised learning. Unsupervised learning uses techniques like clustering and association and supervised learning uses techniques like classification and regression. Now supervised learning is used when you have labeled data. You have historical data then you use supervised learning. When you don't have labeled data then you used unsupervised learning. With in supervised learning there are two types of techniques that are used classification and regression based on what is the kind of problem we are solving. Let's say we want to take the data and classify it. It could be binary classification like a zero or a one. An example of classification we have just seen whether the passenger has survived or not survived. It's like a zero or one that is known as binary classification. Regression on the other hand is you need to predict a value. what is known as a continuous value. Classification is for discrete values. Regression is for continuous values. Let's say you want to predict the share price or you want to predict the temperature that will be the what will be the temperature tomorrow. That is where you use regression. Whereas classification are discrete values is will the customer buy the product or will not buy the product. Will you get a promotion or you will not get a promotion? I hope you're getting the idea. Or it could be multiclass classification as well. Let's say you want to build an image classification model. So the image classification model would take an image as an input and classify into multiple classes whether this image is of a cat or a dog or an elephant or a tiger. So there are multiple classes. So not necessarily binary classification. So that is known as multiclass classification. So we are going to focus on classification because logistic regression is one of the algorithms used for classification.

Now the name may be a little confusing. In fact whenever people come across logistic regression it always causes confusion because the name has regression in it but we are actually using this for performing classification. Okay. So yes, it is logistic regression but it is used for classification. And in case you are wondering is there something similar for regression? Yes, for regression we have linear regression. Keep that in mind. So linear regression is used for regression. Logistic regression is used for classification. So in this video we are going to focus on supervised learning and within supervised learning we're going to focus on classification and then within classification we are going to focus on logistic regression algorithm. So first of all classification. So what are the various algorithms available for performing classification? The first one is decision tree. There are of course multiple algorithms but here we will talk about a few. trees are quite popular and very easy to understand and therefore they used for classification. Then we have K nearest neighbors. This is another algorithm for performing classification. And then there is logistic regression and this is what we are going to focus on in this video and we are going to go into a little bit of details about logistic regression. All right.

What is logistic regression? As I mentioned earlier, logistic regression is an algorithm for performing binary classification. So let's take an example and see how this works. Let's say your car has not been serviced for quite a few years and now you want to find out if it is going to break down in the near future. So this is like a classification problem. Find out whether your car will break down or not. So how are we going to perform this classification? So here's how it looks. If we plot the information along the X and Y axis, X is the number of years since the last service was performed and Y is the probability of your car breaking down. And let's say this information was this data rather was collected from several car users. It's not just your car but several car users. So that is our labeled data. So the data has been collected and um for for the number of years and when the car broke down and what was the probability and that has been plotted along x and y axis. So this provides an idea or from this graph we can find out whether your car will break down or not. We'll see how. So first of all the probability can go from 0 to one. As you all aware probability can be between 0 and one. And as we can imagine it is intuitive as well as the number of years are on the lower side maybe 1 year 2 years or 3 years till after the service the chances of your car breaking down are very limited right so for example chances of your car breaking down or the probability of your car breaking down within 2 years of your last service are.1 probability similarly 3 years is maybe.3 and so on but as the number of years increases. Let's say if it was six or seven years, there is almost a certainty that your car is going to break down. That is what this graph shows. So this is an example of a application of the classification algorithm and we will see in little details how exactly logistic regression is applied here. One more thing needs to be added here is that the dependent variables outcome is discrete. So if we are talking about whether the car is going to break down or not. So that is a discrete value. The y that we are talking about the dependent variable that we are talking about. What we are looking at is whether the car is going to break down or not. Yes or no. That is what we are talking about. So here the outcome is discrete and not a continuous value. So this is how the logistic regression curve looks. Let me explain a little bit what exactly and how exactly we are going to uh determine the class the outcome rather. So for a logistic regression curve a threshold has to be set saying that because this is a probability calculation remember this is a probability calculation and the probability itself will not be zero or one but based on the probability we need to decide what the outcome should be. So there has to be a threshold like for example 0.5 can be the threshold let's say in this case. So any value of the probability below 0.5 is considered to be zero and any value above.5 is considered to be one. So an output of let's say 8 will mean that the car will break down. So that is considered as an output of one and let's say an output of 29 is considered as zero which means that the car will not break down. So that's the way logistic regression works.

Now let's do a quick comparison between logistic regression and linear regression because they both have the term regression in them. So it can cause confusion. So let's try to remove that confusion. So what is linear regression? Linear regression is a process is once again an algorithm for supervised learning. However, here you're going to find a continuous value. You're going to determine a continuous value. It could be the price of a real estate property. It could be your hike, how much hike you're going to get, or it could be a stock price. These are all continuous values. These are not discrete compared to a yes or a no kind of a response that we are looking for in logistic regression. So this is one example of a linear regression. Let's say the HR team of a company tries to find out what should be the salary hike of an employee. So they collect all the details of their existing employees, their ratings and their salary hikes, what has been given and that is the labeled information that is available and the system learns from this. It is trained and it learns from this labeled

Information so that when a new employee's information is fed, based on the rating, it will determine what should be the height. So this is a linear regression problem and a linear regression example. Now, salary is a continuous value. You can get 5,000, 5,500, 5,600. It is not discrete like a cat or a dog or an apple or a banana. These are discrete, or a yes or a no. These are discrete values, right? So this is where you're trying to find continuous values; this is where we use linear regression.

So let's say, just to extend on the scenario, we now want to find out whether this employee is going to get a promotion or not. So we want to find out—that is a discrete problem, right? A yes or no kind of a problem. In this case, we actually cannot use linear regression, even though we may have labeled data. So this is the label data. So based on the employee rating, these are the ratings, and then some people got the promotion, and this is the rating for which people did not get promotion—that is a no. And this is the rating for which people got promotion. We just plotted the data about whether a person has got—an employee has got promotion or not. Yes. No. Right? So there is nothing in between, and what is the employee's rating? Okay, and ratings can be continuous; that is not an issue, but the output is discrete in this case—whether the employee got promotion, yes or no. Okay. So if we try to plot that and we try to find a straight line, this is how it would look, and as you can see, it doesn't look very right because it looks like there will be a lot of errors—this root mean square error, if you remember for linear regression, would be very, very high—and also the values cannot go beyond zero or beyond one. So the graph should probably look somewhat like this, clipped at 0 and 1. But still, the straight line doesn't look right. Therefore, instead of using a linear equation, we need to come up with something different, and therefore the logistic regression model looks somewhat like this.

So we calculate the probability, and if we plot that probability, not in the form of a straight line, but we need to use some other equation. We will see very soon what that equation is. Then it is a gradual process, right? So you see here, people with some of these ratings are not getting any promotions, and then slowly, at a certain rating, they get promotion. So that is a gradual process, and this is how the math behind logistic regression looks. So we are trying to find the odds for a particular event happening, and this is the formula for finding the odd: So the probability of an event happening divided by the probability of the event not happening. So P, if it is the probability of the event happening—probability of the person getting a promotion—and divided by the probability of the person not getting a promotion, that is 1 – P. So this is how you measure the odds. Now, the values of the odds range from 0 to infinity. So when this probability is zero, then the odds will—the value of the odds is equal to zero. And when the probability becomes 1, then the value of the odds is 1/0, that will be infinity. But the probability itself remains between 0 and 1.

Now, this is how an equation of a straight line looks: So y = β0 + β1x, where β0 is the y-intercept and β1 is the slope of the line. If we take the odds equation and take a log of both sides, then this would look somewhat like this. And the term logistic is actually derived from the fact that we are doing this. We take a log of px/(1 – px). This is an extension of the calculation of odds that we have seen. Right? And that is equal to β0 + β1x, which is the equation of the straight line. And now, from here, if you want to find out the value of px, we will see we can take the exponential on both sides, and then if we solve that equation, we will get the equation of px like this: px = 1/(1 + e^–(β0 + β1x)), and recall this is nothing but the equation of the line which is equal to y = β0 + β1x. So that this is the equation also known as the sigmoid function, and this is the equation of the logistic regression L. All right. And if this is plotted, this is how the sigmoid curve is obtained.

So let's compare linear and logistic regression—how they are different from each other. Let's go back. So linear regression is solved or used to solve regression problems, and logistic regression is used to solve classification problems. So both are called regression, but linear regression is used for solving regression problems where we predict continuous values. Whereas logistic regression is used for solving classification problems where we have to predict discrete values. The response variables in case of linear regression are continuous in nature. Whereas here they are categorical or discrete in nature. And the linear regression helps to estimate the dependent variable when there is a change in the independent variable. Whereas here, in case of logistic regression, it helps to calculate the probability or the possibility of a particular event happening. And linear regression, as the name suggests, is a straight line. That's why it's called linear regression. Whereas logistic regression is a sigmoid function, and the curve—the shape of the curve—is S. It's an S-shaped curve.

This is another example of application of logistic regression in weather prediction—whether it's going to rain or not rain. Now, keep in mind, both are used in weather prediction. If we want to find the discrete values like whether it's going to rain or not rain—that is a classification problem—we use logistic regression. But if you want to determine what is going to be the temperature tomorrow, then we use linear regression. So just keep in mind that in weather prediction, we actually use both. But these are some examples of logistic regression. So we want to find out whether it's going to be rain or not, it's going to be sunny or not, it's going to snow or not. These are all logistic regression examples.

A few more examples: Classification of objects. This is a again another example of logistic regression. Now, here, of course, one distinction is that these are multiclass classification. So logistic regression is not used in its original form, but it is used in a slightly different form. So we say whether it is a dog or not a dog. I hope you understand. So instead of saying, is it a dog or a cat or elephant, we convert this into saying—so because we need to keep it to binary classification. So we say, is it a dog or not a dog? Is it a cat or not a cat? So that's the way logistic regression can be used for classifying objects. Otherwise, there are other techniques which can be used for performing multiclass classification. In healthcare, logistic regression is used to find the survival rate of a patient. So they take multiple parameters like trauma score and age and so on and so forth, and they try to predict the rate of survival. All right.

Now, finally, let's take an example and see how we can apply logistic regression to predict the number that is shown in the image. So this is actually a live demo. I will take you into Jupyter notebook and show you the code. But before that, let me take you through a couple of slides to explain what we're trying to do. So let's say you have an 8x8 image, and the image has a number 1, 2, 3, 4, and you need to train your model to predict what this number is. So how do we do this? So the first thing is obviously, in any machine learning process, you train your model. So in this case, we are using logistic regression. So and then we provide a training set to train the model, and then we test how accurate our model is with the test data, which means that, like any machine learning process, we split our initial data into two parts—training set and test set. With the training set, we train our model, and then with the test set, we test the model till we get good accuracy, and then we use it for inference. Right? So that is the typical methodology of training, testing, and then deploying of machine learning models.

So let's take a look at the code and see what we are doing. So I'll not go line by line but just take you through some of the blocks. So the first thing we do is import all the libraries, and then we basically take a look at the images and see what is the total number of images. We can display using matplotlib some of the images, or a sample of these images, and then we split the data into training and test, as I mentioned earlier, and we can do some exploratory analysis, and then we build our model. We train our model with the training set, and then we test it with our test set and find out how accurate our model is using the confusion matrix, the heat map, and use heat map for visualizing this, and I will show you in the code what exactly is the confusion matrix and how it can be used for finding the accuracy. In our example, we got—we get an accuracy of about 94%, which is pretty good. All right, so what is the confusion matrix? This is an example of a confusion matrix, and this is used for identifying the accuracy of a classification model or like a logistic regression model. So the most important part in a confusion matrix is that, first of all, this—as you can see, this is a matrix, and the size of the matrix depends on how many outputs we are expecting, right? So the most important part here is that the model will be most accurate when we have the maximum numbers in its diagonal—like in this case—that's why it has almost 93-94% because the diagonal should have the maximum numbers, and the others—other than diagonals—the cells other than the diagonals should have very few numbers. So here that's what is happening. So there is a 2 here; there are—there is a 1 here. But most of them are along the diagonal. What does this mean? This means that the number that has been fed is 0, and the number that has been detected is also 0. So the predicted value and the actual value are the same. So along the diagonals, that is true. Which means that—let's take this diagonal, right—if the maximum number is here, that means that—like here in this case, it is 34, which means that 34 of the images that have been fed—or rather, actually there are two misclassifications in there. So 36 images have been fed which have number 4, and out of which 34 have been predicted correctly as number 4, and one has been predicted as number 8 and another one has been predicted as number 9. So these are two misclassifications. Okay. So that is the meaning of saying that the maximum number should be in the diagonal. So if you have all of them—so for an ideal model which has, let's say, 100% accuracy, everything will be only in the diagram. There will be no numbers other than zero in all other cells. So that is like a 100% accurate model. Okay. So that's the gist of how to use this matrix—how to use this confusion matrix. So I know the name is a little funny sounding—confusion matrix—but actually it is not very confusing. It's very straightforward. So you are just plotting what has been predicted and what is the labeled information, or what is the actual data—that's also known as the ground truth sometimes. Okay, these are some fancy terms that are used. So predicted label and the actual label—that's all it is. Okay. Yeah. So we are showing a little bit more information here. So 38 have been predicted, and here you will see that all of them have been predicted correctly. There have been 38 zeros, and the predicted value and the actual value is exactly the same. Whereas in this case, right, it has—there are, I think, 37 + 5—yeah, 42 have been fed—the images—42 images are of digit 3, and the accuracy is only 37 of them have been accurately predicted. Three of them have been predicted as number 7, and two of them have been predicted as number 8, and so on and so forth. Okay. All right. So with that, let's go into Jupyter notebook and see how the code looks.

So this is the code in Jupyter notebook for logistic regression. In this particular demo, what we are going to do is train our model to recognize digits, which are the images which have digits from, let's say, 0 to 5 or 0 to 9, and then we will see how well it is trained and whether it is able to predict these numbers correctly or not. So let's get started. So the first part is, as usual, we are importing some libraries that are required, and the last line in this block is to load the digits. So let's go ahead and run this code. Then here we will visualize the shape of these digits. So we can see here, if we take a look, this is how the shape is—1797x64. These are like 8x8 images. So that's what is reflected in this shape. Now, from here onwards, we are basically—once again—importing some of the libraries that are required, like NumPy and Matplotlib, and we will take a look at some of the sample images that we have loaded. So this one, for example, creates a figure, and then we go ahead and take a few sample images to see how they look. So let me run this code, so that it becomes easy to understand. So these are about five images—sample images that we are looking at—0, 1, 2, 3, 4. So this is how the images—this is how the data is. Okay. And based on this, we will actually train our logistic regression model, and then we will test it and see how well it is able to recognize. So the way it works is the pixel information. So as you can see here, this is an 8x8 pixel kind of an image, and the each pixel—whether it is activated or not activated—that is the information available for each pixel. Now, based on the pattern of this activation and non-activation of the various pixels, this will be identified as a 0, for example, right? Similarly, as you can see, so overall, each of these numbers actually has a different pattern of the pixel activation, and that's pretty much that our model needs to learn—for which number, what is the pattern of the activation of the pixels, right? So that is what we are going to train our model. Okay. So the first thing we need to do is to split our data into training and test data set. Right? So whenever we perform any training, we split the data into training and test so that the training data set is used to train the system. So we pass this probably multiple times. And then we test it with the test data set. And the split is usually in the form of—and there are various ways in which you can split this data. It is up to the individual preferences. In our case here, we are splitting in the form of 23 and 77. So when we say test_size as 0.23, that means 23% of the entire data is used for testing, and the remaining 77% is used for training. So there is a readily available function which is called train_test_split. So we don't have to write any special code for the splitting. It will automatically split the data based on the proportion that we give here, which is test_size. So we just give the test_size; automatically, training size will be determined, and we pass the data that we want to split, and the results will be stored in x_train and y_train for the training data set. And what is x_train? These are these are the features, right, which is like the independent variable, and y_train is the label, right? So in this case, what happens is we have the input value which is—or the features value which is in x_train, and since this is a labeled data, for each of them—each of the observations—we already have the label information saying whether this digit is a 0 or a 1 or a 2, so that this—this is what will be used for comparison to find out whether the system is able to recognize it correctly or there is an error. For each observation, it will compare with this, right? So this is the label. So the same way, x_train, y_train is for the training data set; x_test, y_test is for the test data set. Okay. So let me go ahead and execute this code as well, and then we can go and check quickly what is the—how many entries are there in each of these. So x_train, the shape is 1383x64, and y_train has 1383 because there is—nothing like the second part is not required here—and then x_test shape we see is 414. So actually there are 414 observations in test and 1383 observations in train. So that's basically what these four lines of code are saying. Okay. Then we import the logistic regression library, which is a part of scikit-learn. So we don't have to implement the logistic regression process itself. We just call these the function, and let me go ahead and execute that so that we have the logistic regression library imported. Now we create an instance of logistic regression. Right? So logreg is an instance of logistic regression, and then we use that for training our model. So let me first execute this code. So these two lines: So the first line basically creates an instance of logistic regression model, and then the second line is where we are passing our data—the training data set, right? This is our—the predictors—and this is our target. We are passing this data set to train our model. All right. So once we do this—in this case, the data is not large, but by and large, the training is what takes usually a lot of time. So we spend in machine learning activities—in machine learning projects—we spend a lot of time for the training part of it. Okay. So here the data set is relatively small, so it was pretty quick. So all right. So now our model has been trained using the training data set, and we want to see how accurate this is. So what we'll do is we will test it out in probably phases. So let me first try out how well this is working for one image. Okay, I will just try it out with one image—my the first entry in my test data set—and see whether it is correctly predicting or not. So and in order to test it—so for training purpose, we use the fit method. There is a method called fit which is for training the model, and once the training is done, if you want to test for a particular value—new input—you use the predict method. Okay. So let's run the predict method, and we pass this particular image, and we see that the shape is—or the prediction is 4. So let's try a few more. Let me see for the next 10. Seems to be fine. So let me just go ahead and test the entire data set. Okay, that's basically what we will do. So now we want to find out how accurately this has performed. So we use the score method to find what is the percentage of accuracy, and we see here that it has performed up to 94% accurate. Okay. So that's on this part. Now what we can also do is we can also see this accuracy using what is known as a confusion matrix. So let us go ahead and try that as well, so that we can also visualize how well this model has done. So let me execute this piece of code, which will basically import some of the libraries that are required, and we basically create a confusion matrix—an instance of confusion matrix—by running confusion_matrix and passing these values. So we have—so this confusion_matrix method takes two parameters: one is the y_test and the other is the prediction. So what is the y_test? These are the labeled values which we already know for the test data set, and predictions are what the system has predicted for the test data set. Okay. So this is known to us, and this is what the system—the model—

has generated. So we kind of create the confusion matrix, and we will print it. And uh, this is how the confusion matrix looks. As the name suggests, it is a matrix. And um, the key point out here is that the accuracy of the model is determined by how many numbers are there in the diagonal. The more the numbers in the diagonal, the better the accuracy is. Okay.

And first of all, the total sum of all the numbers in this whole matrix is equal to the number of observations in the test data set. That is the first thing, right? So if you add up all these numbers, that will be equal to the number of observations in the test data set, and then out of that, the maximum number of them should be in the diagonal. That means the accuracy is pretty good. If the numbers in the diagonal are less, and in all other places there are a lot of numbers, which means the accuracy is very low. The diagonal indicates a correct prediction. This means that the actual value is the same as the predicted value. Here again, the actual value is the same as the predicted value, and so on. Right? So the moment you see a number here, that means the actual value is something, and the predicted value is something else. Right? Similarly, here the actual value is something, and the predicted value is something else. So that is basically how we read the confusion matrix.

Now, how do we find the accuracy? You can actually add up the total values in the diagonal. So it it's like 38 + 44 + 43 and so on, and divide that by the total number of test observations; that will give you the percentage accuracy using a confusion matrix.

Now let us visualize this confusion matrix in a slightly more sophisticated way, uh, using a heat map. So we will create a heat map with some; we'll add some colors as well. It's uh, it's like a more visually, visually more appealing. So that's the whole idea. So if we, let me run this piece of code, and this is how the heat map uh looks, and as you can see here, the diagonals again uh are all the values are here; most of the values, so which means reasonably this seems to be reasonably accurate, and yeah, basically the accuracy score is 94%. This is calculated as I mentioned by adding all these numbers divided by the total test values or the total number of observations in the test data set. Okay, so this is the confusion matrix for logistic regression.

All right, so now that we have seen the confusion matrix, let's take a quick sample and see how well uh the system has classified, and we will take a few examples of the data. So if we see here, we we picked up randomly a few of them. So this is uh number four, which is the actual value, and also the predicted value; both are four. This is an image of zero. So the predicted value is also zero. The actual value is, of course, zero. Then this is the image of nine. So this has also been predicted correctly; 9, and the actual value is 9, and this is the image of one, and again, this has been predicted correctly as like the actual value. Okay. So this was a quick demo of logistic regression. How to use logistic regression to identify images.

So we put them side by side. We have our linear regression, which is a predictive number used to predict a dependent output variable based on an independent input variable. Accuracy is measured uh using least squares estimation. So that's where you take; you could also use absolute value. Uh, the least squares is more popular. There's reasons for that mathematically and also for computer runtime. Uh, but it does give you an accuracy based on the least square estimation. The best fit line is a straight line, and clearly that's not always used in all the regression models. There's a lot of variations on that. The output is a predicted integer value. Again, this is what we're talking about when we talk about linear regression. And we're talking about regression; it means the numbers coming out. Linear usually means we're looking for that line versus a different model. And it's used in business domain forecasting stocks. Uh, it's used as a basis of most um uh predictions with numbers. So if you're looking at a lot of numbers, you're probably looking at uh a linear regression model.

Uh, for instance, if you do just the high lows of the stock exchange, and you're you're going to take a lot more of that if you want to make money off the stock, you'll find that the linear regression model fits uh probably better than almost any of the other models, even, you know, high-end neural networks and all these other different machine learning and AI models because they're numbers. They're just a straight set of numbers. You have a high value, a low value, volume, uh, that kind of thing. So when you're looking at something as straight numbers um and are connected in that way, usually you're talking about a linear regression model, and that's where you want to start.

A logistic regression model is used to classify a dependent output variable based on an independent input variable. So just like the linear regression model and like all of our machine learning tools, you have your features coming in. Uh, and so in this case, you might have uh label, you know, an image or something like that is is probably the very popular thing right now. Labeling broccoli and vegetables or whatever. Accuracy is measured using maximum likelihood estimation. The best fit is given by a curve. And we saw that um we're talking about linear regression; you definitely are talking about a straight line, although there are other regression models that don't use straight lines. And usually when you're looking at a logistic regression, the math, as you saw, was still kind of a ukitian line, but it's now got that sigmoid activation, which turns it into um a heavily weighted curve. And the output is a binary value between zero and one. And it's used for classification. Image processing, as I mentioned, is is what people usually think of. Um, although they use it for classification of um like a window of things. So you could take a window of stock history, and you could cla generate classifications based on that and separate the data that way. If it's going to be that this particular pattern occurs, it's going to be upward trending or downward trending. In fact, a number of stock uh uh traders use that not to tell them how much to bid or what to bid, uh but they use it as to whether it's worth looking at the stock or not, whether the stock's going to go down or go up, and it's just a zero one. Do I care? Do I even want to look at it?

So, let's do a demo so you can get a picture of what this looks like in Python code. Let's predict the price at which insurance should be sold to a particular customer based on their medical history. We will also classify on a mushroom data set to find the poisonous and non-poisonous mushrooms. And when you look at these two datas, the first one uh we're looking at the price. So the price is a number. Um, so let's predict the price at which the insurance should be sold to. And the second one is we're looking at either it's poisonous or it's not poisonous. So, first off, before we begin the demo, I'm in the Anaconda Navigator. In this one, I've loaded the Python 3.6 and using the Jupyter notebook. And you can use Jupyter notebook by itself. Um, you can use the Jupyter Lab, which allows multiple tabs. It's basically the notebook with tabs on it. Uh, but the Jupyter notebook is just fine. And it'll go into uh Google Chrome, which is what I'm using for my Internet Explorer. And from here, we open up new, and you'll see Python 3. And again, this is loaded with Python 3.6. And we're doing the linear versus logic uh regression or logic. You'll see logit um as one of the one of the names that kind of pops up when you do a search on here. Uh, but it is a logic. We're looking at the logistic regression models, and and we'll start with the linear regression uh because it's easy to understand. You draw a line through stuff.

Um, and so in programming, uh we got a lot of stuff to unfold here in our in our uh startup as we preload all of our different parts. And let's go ahead and break this up. We have at the beginning import uh pandas. So this is our data frame. Uh, it's just a way of storing the data. Think of a uh when you talk about a data frame, think of a spreadsheet. You have rows and columns. It's a nice way of viewing the data. And then we have uh we're going to be bringing in our pre-processing labeling coder. I'll show you what that is. Um, when we get down to it, it's easier to see in the data. Uh, but there's some data in here like um sex. It's male or female. So it's not like an actual number. It's either you're one or the other. That kind of stuff ends up being encoded. That's what this label encoder is right here. We have our test split model. If you're going to build a model, uh you do not want to use all the data. You want to use some of the data, and then test it to see how good it is. And if it can't have seen the data you're testing on until you're ready to test it on there and see how good it is. And then we have our uh logistic regression model, our categorical one, and then we have our linear regression model. These are the two; these right here. Let me just um um clear all that. There we go. Uh, these two right here are what this is all about. Logistic versus uh linear. Is it categorical? Are we looking for a true false, or are we looking for um a specific number? And then finally, um usually at the very end, we have to take and just ask how accurate is our model. Did it work? Um, if you're trying to predict something, in this case, we're going to be doing um, uh, insurance costs, uh, how close to the insurance cost does it measure that we expect it to be? You know, if you're an insurance company, you don't want to promise to pay everybody's medical bill and not be able to. And in the case of the mushrooms, you probably want to know just how much at risk you are for following this model uh as to far as whether you're going to eat a poisonous mushroom and die or not. Um, so we'll look at both of those, and we'll get talk a little bit more about the shortcomings and the um uh value of these different processes.

So let's go ahead and run this. This is loaded the data set on here. And then because we're in Jupyter notebook, I don't have to put the print on there. We just do data set and by, and it prints out all the different data on here. And you can see here for our insurance because that's what we're starting with. Uh, we're loading that with our pandas, and it prints it in a nice format where you can see the age, sex, uh body mass index, number of children, smoker. So this might be something that the insurance company gets from the doctor. It says, "Hey, we're going to; this is what we need to know to give you a quote for what we're going to charge you for your insurance." And you can see that it has uh 1,338 rows and seven columns. You can count the columns. 1 2 3 4 5 6 7. So, there's seven columns on here. And the column we're really interested in is charges. Um, I want to know what the charges are going to be. What can I expect? Not a very good arrow drawn. Um, what to expect them to charge on there. Uh, so is this going to be, you know, is this person going to cost me uh $16,884, or is this person only going to cost me uh 3,866? How do we guess that so that we can guess what the minimal charge is for their insurance?

And then there's one other thing you really need to notice on this data. Um, and I mentioned it before, but I'm going to mention it again because pre-processing data is so much of the work in data science. Um, sex. Well, how do you how do you deal with female versus male? Um, are you a smoker? Yes or no? What does that mean? Region. How do you look at region? It's not a number. How do you draw a line between southwest and northwest? um, you know, they're they're objects. It's either your southwest or your northwest. It's not like I'm southwest. I guess you could do longitudinal latitude, but the data doesn't come in like that. It comes in as true, false, or whatever. You know, it's either your southwest or your northwest. So, we need to do a little bit of pre-processing of the data on here to make this work. Oops, there we go. Okay.

So, let's take a look and see what we're doing with pre-processing. And again, this is really where you spend a lot of time with data science is trying to understand how and why you need to do that. And so we're going to do uh you'll see right up here label uh and then we're going to do the do a label encoder, one of the modules we brought in. So this is sklearns uh label encoder. I like the fact that it's all pretty much automated. Uh, but if you're doing a lot of work with label encoder, you should start to understand how that fits. Um, and then we have uh label.fit right here where we're going to go ahead and do the data set uh sex.drop duplicates. And then for data set sex, we're going to do the label transform the data sex. And so we're looking right here at um male or female. And so it usually just converts it to a zero one because there's only two choices on here. Same thing with the smoker. It's zero or one. So, we're going to transfer the trans uh change the smoker uh 01 on this. And then finally, we did region down here. Region does it a little bit different. We'll take a look at that. And um, it it's I think in this case, it's probably going to do it because we did it on this label transform. Um, with this particular setup, it gives each region a number like 0, 1, 2, 3. So, let's go and take a look and see what that looks like. Go and run this. And you can see that our new data set um has age, that's still a number. Uh, sex is zero or one. Uh, so zero is female, one is male. Number of children, we left that alone. Uh, smoker, one or zero. It says no or yes on there. We actually just do one for no, zero or no. Yeah, one for no. I'm not sure how it organized them, but it turns the smoker into zero or one. Yes or no. Uh, and then region. It did this as uh 0, 1, 2, 3. So there are three regions. Now, a lot of times in in when you're working with data science and you're dealing with uh regions or even word analysis, um instead of doing one column and labeling it 0, 1, 2, 3, a lot of times you increase your features. And so you would have region northwest would be one column, yes or no. Region southwest would be one column, yes or no. True. 01. Uh, but for this this this particular setup, this will work just fine on here.

Now that we spent all that time getting it set up, uh, here's the fun part. Uh, here's the part where we're actually using or set up on this. And you'll see right here we have our um y linear regression uh data set. Drop the charges because that's what we want to predict. And so our x, I'm sorry, our x linear data set, drop the charges because that's what we're going to predict. We're predicting charges right here. So we don't want that as our input for our features. And our y output is charges. That's what we want to guess. We want to guess what the charges are. And then what we talked about earlier is we don't want to do all the data at once. So we're going to take uh three means 30%. We're going to take 30% of our data, and it's going to be as the train as the testing site. So here's our y test and our x test down there. Um, and so that part our model will never see it until we're ready to test to see how good it is. And then of course, right here you'll see our um training set, and this is what we're going to train it. We're going to trade it on 70% of the data. And then finally, the big ones. Uh, this is where all the magic happens. This is where we're going to create our uh magic setup. And that is right here. Our linear model. We're going to set it equal to the linear regression model. And then we're going to fit the data on here. And then at this point, I always like to pull up um if you if you if you're working with a new model, it's good to see where it comes from. And this comes from the scikite uh learn. And this is the sklearn linear model linear regression that we imported earlier. And you can see they have different parameters. The basic parameter works great if you're dealing with just numbers. uh, mentioned that earlier with stock high lows. This model will do as good as any other model out there for do if you're doing just the very basic high lows and looking for a linear fit, a regression model fit. Um, and when you one of the things when I am looking at this is I look for methods, and you'll see here's our fit that we're using right now, and here's our predict. And we'll actually do a little bit in the middle here as far as looking at some of the uh parameters hidden behind it. The math that we talked about earlier. And so we go in this and we go ahead and run this. You'll see it loads the linear regression model and just has a nice output that says, hey, I loaded the linear regression model. And then the second part is we did the fit. And so this model is now trained. Our linear model is now trained on the training data. And so one of the things we can look at is the um um for idx and column name and enumerate xlinear train columns kind of an interesting thing. This prints out the coefficients. Uh, so when you're looking at the back end of the data, you remember we had that formula uh bx1 + bxx2 + the in + the uh intercept uh and so forth. These are the actual coefficients that are in here. This is what it's actually multiplying these numbers by. And you can see like region gets a minus value. So when it adds it up, I guess region, you can read a lot into these numbers. Uh, it gets very complicated. There's ways to mess with them. If you're doing a basic linear regression model, usually don't look at them too closely. Uh, but you might start looking in these and saying, "Hey, you know what? Uh, smoker, look how smoker impacts the cost. Um, it's just massive." Uh, so this is a flag that hey, the value of the smoker really affects this model, and then you can see here where the body mass index uh so somebody who is overweight is probably less healthy and more likely to have cost money, and then of course, age is a factor. Um, and then you can see down here we have uh sexist than a factor also. And it just it changes as you go in there. Negative number. It probably has its own meaning on there. Again, it gets really complicated when you dig into the um workings and how the linear model works on that. And so um we can also look at the intercept. This is just kind of fun. Um, so it starts at this negative number and then adds all these numbers to it. That's all that means. That's our intercept on there. And that fits the data we have on that. And so you can see right here, we can go back and oops, give me just a second. There we go. We can go ahead and predict the unknown data, and we can print that out. And if you're going to create a model to predict something, uh, we'll go ahead and predict it. Here's our Y prediction value linear model predict. And then we'll go ahead and create a new data frame. In this case, from our Xlinear test group. We'll go ahead and put the cost back into this data frame. And then the predicted cost. We're going to make that equal to our y prediction. And so when we pull this up, uh you can see here that we have uh the actual cost and what we predicted the cost is going to be. There are a lot of ways to measure the accuracy on there. Uh, but we're going to go ahead and jump into our

Mushroom Data.

And so, in this, you can see here we we've run our basic model. We built our coefficients. You can see the intercept, the back end. You can see how we're generating a number here. Uh, now with mushrooms, we want a yes or no. We want to know whether we can eat them or not. And so, here's our mushroom file. We're going to go and run this. Take a look at the data. And again, you can ask for a copy of this file. Uh, send a note over to simplylearn.com. And you can see here that we have a class, um, the cap shape, cap surface, and so forth. So there's a lot of features. In fact, there's 23 different columns in here going across. And when you look at this, um, I'm not even sure what these particular like P, E, P, I, don't even know what the class is on this. I'm going to guess by the notes that the class is uh, poisonous or edible.

So, if you remember before, we had to do a little precoding on our data. Uh, same thing with here, uh, we have our cap shape which is B or X or K. Um, we have cap color. Uh, these really aren't numbers. So it's really hard to do anything with just a a single number. So we need to go ahead and turn those into a label encoder, which again there's a lot of different encoders. Uh, with this particular label encoder, it's just switching it to 0, 1, 2, 3 and giving it an integer value. In fact, if you look at all the columns, all of our columns are labels. And so, we're just going to go ahead and uh, loop through all the columns in the data and we're going to transform it into a um label encoder. And so when we run this, you can see how this gets shifted from uh, xbxx to 0, 1, 2, 3, 4, 5 or whatever it is. Class is 0, 1, one being poisonous. Zero looks like it's edible and so forth on here. So we're just encoding it. If you were doing this project, depending on the results, you might encode it differently. Like I mentioned earlier, you might actually increase the number of features as opposed to labeling at 0, 1, 2, 3, 4, 5. Um, in this particular example, it's not going to make that big of a difference how we encode it.

And then, of course, we're looking for uh, the class whether it's poisonous or edible. So we're going to drop the class in our X uh logistics model and we're going to create our Y logistics model is based on that class. So, here's our XY. And just like we did before, we're going to go ahead and split it. Uh, using 30% for test, 70% to program the model on here. And that's right here. Whoops, there we go. There's our uh uh train and test. And then you'll see here on this next setup, um, this is where we create our model. All the magic happens right here. Uh, we go ahead and create a logistics model. I've upped the max iterations. If you don't change this for this particular problem, you'll get a warning that says this has not converged. Um, because then that that's what it does is it goes through the math and it goes, hey, can we minimize the error? And it keeps finding a lower and lower error and it still is changing that number. So that means it hasn't conversed yet. It hasn't find the lowest amount of error it can. And the default is 100. Uh, there's a lot of settings in here. So when we go in here to let me pull that up from the sklearn. Uh, so we pull that up from the sklearn model. You can see here we have our logistic. It has our different settings on here that you can mess with. Most of these work pretty solid on this particular setup. So you don't usually mess a lot. Usually I find myself adjusting the um iteration and I'll get that warning and then increase the iteration on there.

And just like the other model, you can go just like you did with the other model. We can scroll down here and look for our methods. And you can see there's a lot of methods uh available on here. And certainly there's a lot of different things you can do with it. Uh, but the most basic thing we do is we fit our model, make sure it's set right, uh, and then we actually predict something with it. So, those are the two main things we're going to be looking at on this model is fitting and predicting. There's a lot of cool things you can do that are more advanced. Uh, but for the most part, these are the two which um I use when I'm going into one of these models and setting them up. So, let's go ahead and close out of our sklearn setup on there. And we'll go ahead and run this. And you can see here it's now loaded this up there. We now have a uh uh logistic model. And we've gone ahead and done a predict here also just like I was showing you earlier. Uh, so here's where we're actually predicting the data. So we we've done our first two lines of code as we create the model. We fit the model to our training data and then we go ahead and predict for our test data.

Now in the previous model, we didn't dive into the test score. Um, I think I just showed you a graph and we can go in there and there's a lot of tools to do this. We're going to look at the uh model score on this one. And let me just go ahead and run the model score. And it says that it's pretty accurate. We're getting a roughly 95% accuracy. Well, that's good. One 95% accuracy. 95% accuracy might be good for a lot of things, but when you look at something as far as whether you're going to pick a mushroom on the side of the trail and eat it, we might want to look at the confusion matrix. And for that, we're going to put in our y logistic test, the actual values of edible and unedible. And we're going to put in our prediction value. And if you remember on here, uh, let's see. I believe it's poisonous was one. Uh, zero is edible. So, let's go ahead and run that. 0, 1, zero is good. So, here is um a confusion matrix. And this is if you're not familiar with these, we have true, true, false, true, false. So it says out of the edible mushrooms, we correctly labeled 1,211 mushrooms edible that were edible. And we correctly measured 1,113 poisonous mushrooms as poisonous. But here's the kicker. I labeled uh 56 edible mushrooms as being um poisonous. Well, that's not too big of a deal. We just don't eat them. But I measured 68 mushrooms as being edible that were poisonous. So, probably not the best choice to use this model to predict whether you're going to eat a mushroom or not. And you'd want to dig a little deeper before you uh start eating mushrooms off the side of the trail.

So, a little warning there when you're looking at any of these data models, looking at the error and how that error fits in with what domain you're in. Domain in this case being edible mushrooms. Uh, be a little careful. Make sure that you're looking at them correctly. So, we've looked at uh edible or not edible. We've looked at uh regression model as far as uh the end values. what's going to be the cost and what our predicted cost is so we can start figuring out how much to charge these people for their insurance. And so these really are the fundamentals of data science when you pull them together. Uh, when I say data science, I'm talking about your machine learning code. Classification is probably one of the most widely used tools in machine learning in today's world. It is also one of the simpler versions to start understanding how a lot of machine learning works. We're going to start by taking a look at what exactly is classification, the important terminologies around classification. We'll look at some real-world applications, my favorite popular classification algorithms, and there are a lot out there. So, we're only going to touch briefly on a variety of them so you can see how they the different flavors work. And we'll have some hands-on demos in Python embedded throughout the tutorial.

Classification.

Classification is a task that requires the use of machine learning algorithms to learn how to assign a class label to a given data. You can see in this diagram we have our unclassified data. It goes through a classification algorithms and then you have classified data. It's hard to just see it as data and that really is where you kind of start and where you end when you start running these u machine learning algorithms and classification. And the classification algorithms is a little black box in a lot of respects. And we'll look into that. You can see what I'm talking about when we start swapping in and out different models. Let's say we are given the task of classifying a given bunch of fruits and vegetables on the basis of their category. I.e. fruits are to be grouped together and vegetables are to be grouped together. And so we have a data set. We'll call it the bunch is divided into clusters. One of which consists of the fruits while the other has the vegetables. You can actually look at this as any kind of data. When we talk about breast cancer, can we sort out an images to see what is malignant, what is benign, very popular one. Can you classify flowers? The iris uh data set. Uh, certainly in wildlife, can you classify different animals and track where they're going? Classification is really the bottom starting point or the baseline for a lot of machine learning and setting it up and trying to figure out how we're going to break the data up. So we can use it in a way that is beneficial. So here the fruits and the vegetables are grouped into clusters and each clusters has a specific characteristic i.e. whether they are a fruit or a vegetable and you can see we have a pile of fruits and vegetables. We feed it into the algorithm and the algorithm separates them out and you have fruits and vegetables.

So some important terminologies before we dig into how it sorts them out and what that all means. Uh, when we look at the terminologies you have a classifier that's the algorithm that is used to map the input data to a specific category, the classification model, the model that predicts or draws a class to the input data given for training, feature, it is an individual measurable property of the phenomena being observed and labels the characteristics on which the data points of a data set are categorized, the classifier and the classification model go together a lot. Lot of times the classifier is part of the classification model. Um, and then you choose which classifier you use after you choose which model you're using. Where features are what goes in, labels are what comes out. So your classifier models right in the middle of that. That's that little black box we were just talking about. Clusters, they are a group of data points which have some common characteristics. Binary classification. It is a classification condition with two outcomes which are either true or false. Multilabel classification. This is a classification condition where each sample is assigned to a set of labels or targets. Multiclass classification. The classification with more than two classes. Here each sample is assigned to one and only one label.

When we look at this group of terminologies, a few important things to notice. Uh, going from the top clusters. When we cluster data together, we don't necessarily have to have an end goal. We just want to know what features cluster together. These features then are mapped to the outcome we want. In many cases, the first step might not even care about the outcome, only about what data connects with other data. And there's a lot of clustering algorithms out there that do just the clustering part. Binary classification, it is a classification condition with two outcomes which are either true or false. we're talking usually um it's a uh it's either a cat or it's not a cat. It's either a dog or it's not a dog. Um, that's the kind of thing we talk about binary classification. And then that goes into multi-lel classification. Think of label as you can have an object that is brown. You can have an object that is labeled as a dog. So it has a number of different labels. That's very different than a multiclass classification where each one's a binary. uh, you can either be a cat or a dog. You can't be both a cat and a dog.

Real-world applications.

So to make sense of this uh of course the challenge is always in the details is to understand how we apply this in the real world. So in real-world applications we use this all the time. We have email spam classifier. So you have your uh email inbox coming in. uh, it goes through the email filter that we usually don't see in the background and it goes this is either valid email or it's a spam and it puts it in the spam filter if that's what it thinks it is. Uh, Alexa's voice classifier, Google voice, any of the voice classifiers, they're looking for points. So they try to group words together and then they try to find those groups of words trigger a class a classifier. So it might be that the classifier is to open your tasks program or open your text program so that you can start sending a text. Sentimental sentiment analysis is really big. Uh, when we're tracking products, we're tracking marketing. Trying to understand uh whether something is liked or disliked is huge. Uh, that's like one of the biggest driving forces in sales nowadays. And you almost have to have these different filters going on if you're running a large business of any kind. Fraud detection. Uh, you can think of banks. Uh, they find different things on your bank statement and they detect that there's something going on there. They have algorithms for tracking the logs on computers. They start finding weird logs on computers. They might find a hacker. I mentioned the cat and dog. So here's our image classification. We have a neighbor who runs an outdoor webcam and we like to have it come up with a classification when the wild animals in our area are out like foxes. We actually have a mountain lion that lives in the area. So it's nice to know when he's here. Handwriting prediction, uh, classifying A, B, C, D, and then classifying words to go with that. So, let's go ahead and, uh, roll our sleeves up and take a look at some popular classification algorithms.

Before we look at the algorithms, uh, let's go back and take a look at our definitions. We have a classifier and a classification model. So, we're looking at the classifier, an algorithm that is used to map the input data to a specific category. One of those algorithms is a logistic regression. The logistic regression is a classification algorithm used to model the probability of a certain class or event existing such as pass fail or win lose etc. It provides its output using the logistic function or sigmoid function to return the probability value that can then be mapped to two or more discrete classes. A sigmoid function is an activation function that fits the variable and limit the output to a range between zero and one. A standard sigmoid function or logistic function is represented by the formula fx = 1 / 1 + e the minus x where x is the equation of the line and e is the exponential. Just taking a quick look at this um you can think of this as being a point of uncertainty. And so as we get closer and closer to the middle of the line, it's either um activated or not. And we want to make that just shoot way up. Uh, so you'll see a lot of the activation formulas kind of have this nice scurve where it approaches one and approaches zero. And based on that, there's only a small region of error. And so you can see in the sigmoid logistic function uh the 1 over 1 + e minus x to the minus x. You can see it frames that that nice scurve. Uh, we also can use a tangent variation. There's a lot of other different uh models here as far as the actual algorithm. This is the most commonly used one. Let's go ahead and roll up our sleeves and take a look at a demo that is going to use a logistic regression. So we're going to have the activation formula and the model because you have to have you have to have both.

For this we will go into our Jupyter notebook. Now, I personally use the Anaconda Navigator to open up the Jupyter notebook to set up my IDE as a web-based. It's got some advantages that it's very easy to display in uh but it also has some disadvantages in that if you're trying to do multi-threads and multiprocessing, you start running into a single git issues with Python and then I jump to uh PyCharm. Really depends on whatever ID you want. Just make sure you've installed uh Numpy and the sklearn modules into your Python in whatever environment you're working in so that you'll have access to that for this demo. Now, the team in the back has prepared my code for me which I'll start bringing in one section at a time so we can go through it. Before we do that, it's always nice to actually see where this information is coming from and what we're working with. Uh, so the first part is we're going to import our packages which you need to install into your Python and that's going to be your numpy. Um, we usually use numpy as inp. And then from sklearn the learn model we're going to use a logistic regression. And from sklearnmetrics we're going to import the classification report confusion matrix. And if we go ahead and open up s uh the scikit-learn.org or and go under their API, you can see all the different um features and models they have. And we're looking at the linear model uh logistic regression, one of the more common classifiers out there. And if we go ahead and go into that and dig a little deeper, you'll see here where they have the different settings. And it even says right here, note the regularization is applied by default. So by default, that is the activation formula being used. Now, we're not going to spend we might come back to this look at some of the other models because it's always good to see what you're working with. But let's go ahead and jump back in here. And we have our imports. We're going to go ahead and run those. Uh, so these are now available to us as we go through our Jupyter notebook script. And they put together a little piece of data for us. This is simply um going through uh 0 to 1. Actually, let's go ahead and print this out over here. We'll go ahead and print x just so you can see we're actually looking for. And when we run this, you can see that we have our x is 0, 1, 2 through 9. We reshaped it. The reason for this is just looking for a row of data. Usually we have multiple features. We just have the one feature which happens to be 0 through nine. And then we have our 10 answers right down here. Uh, 0, 1, 0, 0, 1, 1. You can bring in a lot of different data depending on what you're working with. Uh, you can make your own. You can instead of having this as just a single, you could actually have like multiple features in here, but we just have the one feature for uh this particular demo. And this is really where all the magic happens right here. Uh, and I told you it's like a black box. That's the part that is is is kind of hard to follow. And so if you look right here, we have our model. We talked about the model right there. And then we went ahead and set it for library linear. As I showed you earlier, that's actually default. So it's not that important. uh, random state equals zero. This stuff you don't worry too much about. And then with the scikit learn, you'll see the model fit. This is very common to scikit. They use similar stuff in a lot of other different packages, but you'll you'll see that that's very common. You have to fit your data. And that means we're just taking the data and we're fitting our X right here, which is our features. That's our X. And here's Y. These are the labels we're looking for. So before we were looking at is it fraud, is it not, is it cat, is it not um that kind of thing. And this is looking at zero. So we want to have a binary setup on this. And we'll go ahead and run this. Uh, you can see right here, it just tells us what we loaded it with as our defaults. And that this model has now been created and we've now fit our data to it. And then comes the fun part. You work really hard to clean your data to um bake it and cook it. There's all kinds of I don't know why they go with the cooking terms as far as

How we get this data formatted? Then you go through and you pick your model, you pick your solver, and you have to test it to see, hey, which one's going to be best. And so we want to go ahead and evaluate the model.

And you do this is that once you've figured out which one is going to work the best for you, you want to evaluate it so you can compare it to your last model. And you can either update it to create a new one or maybe change the um solver to something else. I mentioned tangent. That's one of the other common ones that's commonly used with language. For some reason, the tangent, even though it looks almost to me identical to the one we're using with the sigmoid function, uh it for some reason it activates better with language, even though it's a very small shift in the actual math behind it.

We already looked at the um data early, but we'll go and look at it again just so you can see. We look at we have our rows of 01 row. It only has one entity. And we have our output that matches these rows. And these do have to match. You'll get an error if you put in something with a different shape. So, if you have um 10 rows of data and nine answers, it's going to give you an error because you need to have 10 answers for it. A lot of times you separate this too when you're doing larger models. Uh but for this, we're just going to take a quick look at that.

The first thing we want to start looking at is uh the intercept, one of the features inside our linear regression model. We'll go ahead and run that and print it. Uh you'll see here we have an intercept of -1.516. And if we're going to look at the uh intercept, we should also look at our um coefficients. And if you run that, you'll see that we get a list. We get the um our coefficient is the 7035. You can just think of this as your uklitian geometry for very basic model like this where it intercepts the y at some point and we have a coefficient multiplying by it. A little more complicated in the back end, but that is the just this simple model with just the one feature in there.

And we'll go ahead and uh we'll reprint the y because I want to put them on top of each other with the y predict. And so these were the y values we put in. And this is the y predict we had coming out. And you can see um yeah, here we go. Uh there's the y actual and there's what the prediction comes in. Uh now keep in mind that we used the actual complete data as part of our training. Uh that is if you're doing a real model a big stopper right there because you can't really see how good it did unless you split some data off to test it on. This is the first step is you want to see how your model actually test on the data you trained it with. And you can see here there is this point right here where it has it wrong and this point right here where it also has it wrong. And it makes sense because we're going our input is 0 1 0 through 9ine and it has to break it somewhere and this is where the break is. Uh, so it says this half the data is going to be zero because that's what it looked like to me if I was looking at it without an algorithm. And this data is probably going to be one. And I I didn't I forgot to point this out. So let's go back up here. I just kind of glanced over this window here where we did a lot of stuff. Let's go back and and just take a look at that.

What was done here is we ran a prediction. Uh, so this is where our predict comes in is our model.predict. So we had a model fit. We created the model. We programmed it to give us the right answer. Uh now we go ahead and predict what we think it's going to be. There's our model.predict probability of X. And then we have our Y predict which is very similar but this is has to do more with the probability numbers. So if you remember down below we had the setup where we're looking at uh that sigmoid function. That's what this is returning and the Y predict is returning a zero or a one. And then we have our um confusion matrix. We'll look at that. And we have our report which just basically compares our Y to our Y predict which we just did. It's kind of nice. It's simple data. So it's really easy to see what we're doing. That's why we do use a simple data. This can get really complicated when you have a lot of different features and things going on and splits.

Uh so here we go. We've had our I printed out our actual and our prediction. So this is the actual data. This is what the predict ran. Um, and then we'll go ahead and do we're going to print out the confusion matrix. We were just talking about that. Uh, this is great if you have a lot of data to look at. But you can see right here, our confusion matrix says, uh, if you remember from the confusion matrix, we have the two. This is two, correct? One, two. And, uh, it's been a while since I looked at a a confusion matrix. There's the two. And then we have this one, which is our six. That's where the six comes from. And then we have this one which is the U one false. This is a two one. So we have this one here and this one here which is misclassified. This really depends on what data you're working with as to what your is important. Um you might be looking at this model and if this model this confusion matrix comes up and says that uh you've mclassified even one person as being non malignant cancer. That's a bad model. Uh I wouldn't want that classification. I'd want this number to be zero. I wouldn't care if this false positive was off by a little bit more as long as I knew that I was correct on the important factor that I don't have cancer. So you can see that this confusion matrix really aims you in the right direction of what you need to change in your model, how you need to adjust it.

Uh and then there's of course a report. Reports are always nice. Um if you notice we generated a report earlier. We'll go and just print the report up. And you can remember this is our report. It's a classification report. Y predict. So we're just putting in our two values. Basically what we did here visually with our actual and our predicted value. And we'll go ahead and run the report. And you can see it has the precision, uh, the recall, your F1 score, your support, uh, translated into a accuracy, macro average, and weighted average. So, it has all the numbers. A lot of times when working with, um, clients or with the shareholders in the company, this is really where you start because it has a lot of data and they can just kind of stare at it and try to figure it out. And then you start bringing in like the confusion matrix. I almost do this in reverse as to what they show. I would never show your shareholders the intercept or the coefficient. That's for your internal team only working on machine language. Uh but the confusion matrix and the report are very important. Those are the two things you really want to be able to show on these. Uh and you can see here we did a decent job of um classifying the data. Managed to get a significant portion of it correct. uh we had our was it accuracy here is a 080 F1 score uh that kind of thing. So you know it's a pretty accurate model of course this is pretty goofy because it's very simple model and it's just splitting the model between uh ones and zeros.

So what is deep learning? Deep learning is a subset of machine learning which itself is a branch of artificial intelligence. Unlike traditional machine learning models which require manual feature extraction, deep learning models automatically discovers representation from raw data. So this is made possible through neural networks particularly deep neural networks which consist of multiple layers of interconnected nodes. So these neural network are inspired by the structure and the function of human brain. Each layer in the network transform the input data into more abstract and composite representation. For instance, in image recognition, the initial layer might detect simple features like edges and textures while the deeper layer recognizes more complex structure like shapes and objects. So, one of the key advantage of deep learning is its ability to handle large amount of unstructured data such as images, audios and text making it extremely powerful for various application. So stay tuned as we delve deeper into how these neural networks are trained, the types of deep learning models and some exciting applications that are shaping our future types of deep learning. Deep learning AI can be applied supervised, unsupervised and reinforcement machine learning using various methods for each.

The first one supervised machine learning. In supervised learning, the neural network learns to make prediction or classify that data using label data sets. Both input features and target variables are provided and the network learns by minimizing the error between its prediction and the actual targets. A process called back propagation. CNN and RNN are the common deep learning algorithms used for tasks like image classification, sentiment analysis and language translation.

The second one unsupervised machine learning. In unsupervised machine learning, the neural network discovers patterns or cluster in unlabelled data sets without target variables. It identifies hidden pattern or relationship within the data. Algorithms like autoenccoders and generative models are used for tasks such as clustering, dimensionality reduction and anomaly detection.

The third one, reinforcement machine learning. In this, an agent learns to make decision in an environment to maximize a reward signal. The agent takes action, observes the records and learns policies to maximize cumulative rewards over time. Deep reinforced learning algorithms like deep networks and deep deterministic polygradient are used for tasks such as robotics and gameplay.

Moving forward, let's see what are the artificial neural networks. Artificial neural networks inspired by the structure and the function of human neurons consist of interconnected layers of artificial neurals or units. The input layer receives data from the external resources and it passes to one or more hidden layers. Each neuron in these layers computes a weighted sum of inputs and transfers the result to the next layer. During training, the weight of these connection are adjusted to optimize the network's performance. A fully connected artificial neural network includes an input layer or more hidden layers and an output layer. Each neuron in a hidden layer receives input from the previous layer and sends its output to the next layer. So this process continues until the final output layer produce the network response. So moving forward let's see types of neural networks.

So deep learning models can automatically learn feature from data making them ideal to task like image recognition, speech recognition and natural language processing. So the most common architecture and deep learnings are the first one feed forward neural network FN. So these are the simplest type of neural network where information flows linearly from the input to the output. They are widely used for tasks such as image classification, speech recognition and natural language processing NLP.

The second one convolutional neural network designed specifically for image and video recognition. CNN's automatically learn feature from images making them ideal for image classification, object detection and image segmentation.

The third one, recurrent neural networks, RNNs are specialized for processing sequential data, time, series, and natural language. They maintain an internal state to capture information from previous input making them suitable for task such as a speech recognition, NLP, and language transition.

So now let's move forward and see some deep learning application. The first one is autonomous vehicle. Deep learning is changing the development of self-driving car. Algorithms like CNN's process data from sensors and cameras to detect object, recognize traffic signs and make driving decision in real time, enhancing safety and efficiency on the road.

The second one is healthcare diagnostic. Deep learning models are being used to analyze medical images such as X-rays, MRIs and CT scans with high accuracy. They help in early detection and diagnosis of diseases like cancer, improving treatment outcomes and saving lives.

The third one is NLP. Recent advancement in NLP powered by deep learning models like transformers, chat GPD have led to more sophisticated and humanlike text generation, translation and sentiment analysis. So application include virtual assistant, chat bots and automated customer service.

The fourth one defake technology. So deep learning techniques are used to create highly realistic synthetic media known as defects. While this technology has entertainment and creative application, it also raises ethical concern regarding misinformation and digital manipulation.

The fifth one, predictive maintenance in industries like manufacturing and aviation. Deep learning models predict equipment failures before they occur by analyzing sensor data. The proactive approach reduces downtime, lowers maintenance cost, and improves operational efficiency.

So now let's move forward and see some advantages and disadvantages of deep learning. So first one is high computational requirements. So deep learning requires significant data and computational resources for training. Whereas advantage is high accuracy achieves a state-of-the-art performance in tasks like image recognition and natural language processing. Whereas deep learning needs large label data sets often require extensive label data set for training which can be costly and time consuming together. So second advantage of deep learning is automated feature engineering automatically discovers and learn relevant features from data without manual intervention. The third disadvantage is overfitting. So deep learning can overfit to training data leading to poor performance on new unseen data. Whereas the third deep learning advantage is scalability. So deep learning can handle large complex data set and learn from massive amount of data.

So in conclusion, deep learning is a transformative leap in AI mimicking human neural networks. It has changed healthcare, finance, autonomous vehicles and NLP. Yeah. Step by step we will go through all of this and uh we'll make sure that we learn everything and we bring everything together towards the end. Right. Without further ado, let me just straight away deep dive to business, right? To learn data science, right? And with this data science there is also something which is prefixed which is applied data science and suffix for this is with Python right apply data science with Python right so there are there are two key concepts which are going to be a part of this course the first one is the knowledge about data science that what data science is and Then because we are doing an applied course right we are doing an applied course I will try to tie up these concepts which we will understand in data science with a tool right which is Python for you right we already know about uh 60 65% of Python right which is the fundamental Python and now we will be moving to the next step to advanced Python right and using Python leveraging Python we will be solving a lot of problems of data science right using this okay so the first few sessions right the first few sessions will be about making you a breast with Python what Python is right what how and what packages do we have how do they work in reality right and all those things and then we will be coupling it up with the data science concepts and then finally towards the end of the session in the Last few classes we will be doing uh we will be taking a real data set and on that data set we will be applying all these concepts right to understand the data better and we will be drawing inferences from that to convert that into information to take actionable insights using that actionable insights taking a better decision right we'll do all of that in in the actual way okay so now guys if you understand this then the next point of contention is data sets, right? One of the most famous keywords on the planet right now, right? One of the most famous keywords in the planet right now. Do you think that these two things? Okay, let me put it different way. What do you think that can be the possible explanation about this term data science? You know data, you know science. What do you think is going to follow in these sessions? What is data science to you as per these two words?

Okay, so this is people made up of two words, right? Data and science, right? So what we are trying to do is we are trying to understand data, right? We are trying to understand data, right? and then do something to it, right? Understanding its science, understanding the uh nature, the behavior of this data and converting it in something called as information, right? Do do we know difference between data and information? Data is something which is completely raw. Okay? It is completely raw. It has no meaning. Right? It has no meaning, isn't it? For example, I give you these stats of some player like suppose Sid Dhoni. I give you stats of Mahindra Singh Dhoni, right? That what what what were his scores? Uh what is his name? What is his age and you know all those things now everything is there, right? But we don't know what to do about it, right? Do you think the score of Dhoni has any context people? It has any context? No. Right. But when I deep down, but I when I go and deep dive about it, right? What is the first thing you find out of scores? What is the first thing you find out of score scores? You try to find the average of score. Isn't it that in last 10 innings, right? In last 10 innings. Before this also, you need something which is called as a problem statement. Isn't it? Now for example the problem statement is selectors want to understand selectors wants to understand that whether Mahindra Singh Dhoni should be picked up. So there's a problem now right selectors want to see that whether Dhoni is fit for the next tournament or not. So what we will do we will now try to take the mean of the scores right for last 10 innings and if this score is suppose X we will try to compare this with Y. What is Y? Y is a reference right Y is a reference that we want to compare it against. Now when you are doing this comparisons when you are applying these techniques to this score this is now slowly becoming information right and at the end of the day once you have the strike rate once you have the mean score of Dhoni once you have his age once you have his fitness score all those things will now help you to take this particular decision because now what you have is called has information, right? Because this has context, right? This has meaning and this is usually processed, right? This is usually processed, right? This is usually processed. Now, what did we do? Now, what did we do here? If you will go and read about data science, data science says, data science says it is the art of collecting, right? Cleaning, analyzing, modeling, improving Right. And visualizing Right. Visualizing the data. Right? If a person is adept in doing all these things, this person people is cumulatively called a data scientist. Right? That person is called a data scientist. Right? So before going to the definition of data scientist now I will give you some more examples right I'll give you some more examples data science people as I said is a combination of these things right you have to collect the data right you have to collect the data right and this has a lot of things data can be collected from two types in two types one is primary and the second one is secondary Right. What is the primary way of collecting data? From your IoT devices, right? From sensors, from your inbuilt machines, right? Then from your surveys which you float, right? Questionnaires, right? All these things are primary ways. What is the secondary way of collecting data? Purchasing data. Right? Using internet data because you have not generated it. You are just using someone else's data. Right? Something like uh transfer learning. What is transfer learning? Transfer learning is a technique where suppose I am bank A and you are bank B. Right? So bank A has created some model right trained on their data. Now you are going to use the exact same model, right? You're going to use the exact same model, right? Maybe you're not seeing the data but you are just using the property of data like mean, median, mood and a lot of modeling things which will come we will learn about them and you use this model on your particular data. Right? So in a way you did not have enough data to create the model yourself but you are now using someone else's model to run your data on it. Right? So this is called as transfer learning. So this kind of collection is basically secondary data collection. So you can collect the data right then you can perform data analysis right you can perform data analysis right. How will you perform this data analysis using complex algorithms right? Using complex algorithms. Right? some statistics, right? You can use artificial intelligence, right? Artificial intelligence, you can use machine learning, right? Right. You can use

All these things are for data analysis. Then you can transform, transform the patterns into predictions, right? You can transfer these patterns into predictions, right? Which can be used for business decision making, right? For business decision making. Then you can validate the results, right, and present the results, right? So this is like a complete life cycle of a data scientist, right? So before going further, let me give you what combinations you need to have to become a data scientist. The first one is domain knowledge, right.

So what is domain knowledge? First of all, I told you, right, there will be a problem, right? You'll be solving a problem in any project of data science. You'll be trying to solve a problem, right? And the problem will be belonging to a particular domain. Even if you are working for yourself, right? Even if you're an entrepreneur, then also you'll be solving a problem. So this domain knowledge part includes things like understanding—it's a very important diagram—understanding client requirement, right? Understanding the client requirement, right? Important criterians, right? Important criteria knowledge, right? For example, to give an example, suppose we have created a machine learning model. Okay, understanding the data, we have created a machine learning model whose accuracy is 90%, right? Is 90% a good accuracy? Yeah, fairly decent accuracy. Yes. Suppose you have to predict sales, right? You are selling something. Suppose you are selling clothes and you want to predict what will be the sales for the next week. When you use this model, whatever the output model gives you, what is going to be the accuracy of your output using this model? How much accuracy? 90%. But my question is, the model is 90%. But my question to you is that if 90% accuracy is on sales data, a person like me will be very, very happy. Okay? Very, very happy. I'll be probably dancing, right? But if you try to apply the same model, right, for a medical diagnosis case, will you be interested in getting operated in such a hospital or an institution where the accuracy is coming as 90%? Domain knowledge, right? Domain knowledge. We need to understand what are the exact requirements. We need to understand what are the exact expectations, right? And we need to know how much do we need to pivot, right? So the first thing in data science is these accuracies and everything are subjective, right? They are subjective. So for that, you need domain knowledge. Domain knowledge part very important, guys, very important—these three things which I'm going to tell.

Second part is people—the game changer, right? The second part is computer science. Now if I take you back in history, okay, if I take you back in history in the 1980s or somewhere, then do you—okay, how many of you think that data science is a new concept? How many of you think that data science is a new concept? I hope you all—it's not a new concept. Everyone knows that. Yeah, it has been happening for ages. Just like you guys will be shocked if you already don't know—AI was coined in the year 1956, 1956, at the University of Dartmouth. AI was coined by Paul McCarthy, right? And we saw the boom of AI in the year 2010, right? Such a long journey. Same case with data science because people back in the day—data science was called as data mining. Everyone heard about it—data mining, knowledge databases. Yeah, we need—we used to mine the data. Now, what were the problems? What were the hiccups of data mining? The hiccups for data mining was that we were doing everything, everything manually, right? Now if I give you 100 points, can you calculate the mean? Or let's say if I give you two points to multiply—2 * 3, how much time will you take? 2 seconds. Yep. How much time a computer will take? 2 seconds. If I give you to multiply 2,489 * 200, how much time will you take to calculate this? Say 5 seconds. How much time a computer will take? 2 seconds. Now if I give you to multiply 2,489 * 15,195,4386, how much time will you take to calculate this manually? Maybe say 1 minute, 80 seconds, 1 minute. How much time a computer will take? Still 2 seconds, right? Still 2 seconds, right? And if I give you to calculate this over 200 times, you will take 200 minutes, right? Using parallel computing, a computer will still take about 3 to 5 seconds, right? So are you understanding the power of computer? Do you understand this concept in this relationship? What was happening back then? What computer science did—people—was it revolutionized the way data mining was happening, and that thing now is called as—that again—that thing now is called as data science in which computer science is one of the most important contributors. So this is just one reason. Now manually, right? Manually, if I give you say 1 million rows of data, right? 1 million rows of data, right? So how many pages will you—pages will you need to store this data? Suppose your notebook is like this, right? These boxes and here you are storing the data—1 million. So maybe you can buy n number of notebooks, but now do you think it's as easy in storing something in a computer? Because back in the day, people—we had memory issues, isn't it? Memory constraints. So there is something called as Moore's Law, right? Which says as the advancement in microprocessors will increase, the price of microprocessor will decrease, right? So this is what is happening, right? Now, back in the 1980s, if I show you, right, guys, right, yeah, 2.5 kg exactly, right? It was the size of a fridge—hard disk, but now it fits in your palm, right? So this was enabled—the storage techniques, right? The processing techniques, tech—the infrastructure, right? Things like big data—what kind of data do you think we will be dealing with, people, in data science? You all know the term—very famous term—the kind of data—big data, right? Everyone knows about big data. What is big data? Yes, a data which is fast, right? It has velocity, veracity, variety, right? So this kind of data needs to be stored; this kind of data needs to be processed. So which thing brought all these things into data science? It was given to us by computer science, right? So computer science people included things like database management, right? Data validation, right? Data infrastructure, right? Data infrastructure, right? Then we had languages, computer languages—which is Python, right now for us, right? Again, do you think when you do this thing manually, right? Suppose you do this thing manually, how easy do you think it will become using something like Python or any other computer language to create complex models? How easy it will be to do that—to create the complexity in models, right? Where you can capture the nonlinear nature. Isn't it, people? Isn't it? For example, for example, let me tell you this: 2, 4, 6, 8, 10, dash. What do you think is the next number, guys? 12. If I tell you to define this to me in—Okay, let's leave that. What do you think is going to be the next number here? What is the next number? 11. Next number—25. Perfect. Now, guys, if I ask you to write these numbers, right, the way you predicted them, can you give me a function—f(x) is equal to—what is the f(x) here? It's 2x, right? It's 2x. f(x) is equal to 2x. If I tell you to create a function, it will be f(x) is equal to 2x. What will be the function here, guys? f(x) will be equal to x + 1. Yeah, x + 1. Yeah. And here f(x) will be equal to x². Yeah. Now the last example, right? Last example. What is the next number here? I don't want the number. I want the function. I want this so that I can generalize. Isn't it? How did you reach this figure? How many of you think it's not possible to determine this? How many of you think it is not possible to determine this? Yeah. How many of you think—what if I just change this question and ask you—how many of you think it is not possible to determine this manually? Same response, but now I say, how many of you think it is not possible to determine this with computers? Will your answer still remain no? Do you think I cannot approximate this function using computers? We have something called as deep neural networks, and they are called as universal function approximators, right? So this is the problem, people; this is the problem, right? I will show this to you when the time comes, right? I will remember this example and I will show this to you, but now what I'm trying to tell you is the things which seemed impossible manually were solved by—what? It was solved by computers—the distribution of this is like this—can you figure it out yourself? No, right? We cannot, isn't it? We cannot do that. So this kind of approximation will be given by—what? It will be only given by machines, right? And this is, people, what data science is all about, right? It is what computer science did inside data science, right? I hope this is clear. I'm assuming a lot of you will be going for interviews and everything after this course, right? So this will be a very, very important thing for you to know, right? Often it is asked why data science is having computer science in it, right? The reason is this. Okay, so this is the role of computer science inside data science.

Now, people, the third thing—the third circle, which is one of the most parts—is maths and stats, right? Mathematics, statistics, which was optimization, right? Optimization of your models, right? Design of model, right? Now, guys, if you look carefully, if you look carefully, in order to approximate this, what you—what will you be playing with? You will be playing with a lot of data. You will be playing with a lot of mathematical tools, mathematical concepts, and statistical concepts, isn't it? How did you—what do you call this? This is maths, right? This is statistics and mathematics. Finding mean, median, mode, standard deviations, probability, statistics—all these will lead to this kind of result, isn't it? So this becomes the third wheel of this particular—of—of this particular diagram. And this point of intersection, right? The sweet point of intersection is basically data science. Yeah, is particularly data science, right? So now, people, this point—okay, this point is basically representing data engineering, right? Data engineering, right? Data engineering is, people, the part of data science which enables us to capture the correct data, right? How the data will flow, how the data will be stored, how the data will be cleaned, right? All this is done by—home? It is done by data engineering. Just because—so am I right, guys? You'll go for interviews, right, after this and try to fetch yourself jobs in this domain—data science, AI, ML—if my understanding is correct. Is that the aim? Yes. So now, guys, there will be three types of companies, or let's say, to simplify, let's say two types: one is small, and the other one is big, right? So in a small organization, if you become a part of a small organization and you are the data scientist there, you can be involved in all of these things, right? All of these things are possible, right? Your bosses and your management will expect you to construct all these flows, right? Know computer science, you should know maths and stats, and you should have the domain knowledge, and you will be asked to do all of this. But if you are going to become a part of a big organization, usually all these roles are fragmented. All of these roles are fragmented, right? There's a separate data infrastructure team. There's a separate data governance team. Now, guys, when you go on to collect the data, can you collect any sensitive data about—is it possible ethically? It's not right, and legally also it's not right. My question to you is, who will look after this compliance? Whose responsibility indeed it is to look after this compliance? Data scientist. So this is about fragmentation. If you are part of a big organization, this thing will be taken by someone else, right? But if you are a part of a small organization, you will be—know—you'll be expected to do this all by yourself. In—if—if you are part of a big organization, then do you think you need to have this domain knowledge? The answer is no. Why? Because there will be a separate set of people who are called as what? Who are called as business analysts. Have you heard about these positions, people? Business analysts. What is a business analyst role? It is basically a technical translator, right? Who knows technical, who knows domain, and that person goes and talks to the client, talks to the client in a layman language, converted into technical requirement coupled with the domain knowledge and give you the document. This is basically a medical engineering problem. So we need this, this, this, this, and you need to fulfill this, this, this, this criteria. But again, if you're part of a small organization, who needs to take care of that? Who needs to make sure that you know everything about a domain? You yourself, right? You yourself, right? Then, people, this area usually represents whom? This area represents research and analysis. Why? Because they—we have people who have domain knowledge, and we have people who are knowledgeable in math and stats. Have you heard about a position called as actuary in the world? Actuarial science. Actuaries are people who are basically dealing with domains which are very, very heavily data intrinsic, right? For example, finance domain, right? Finance is all about numbers. So there we go—actuarial science, and we have math, stats, optimization, model development, all of those things happening there. We are also, at this point of time, people belonging to with section—research and data analysis, right? And then, people, there's a third intersection, right? There's a third—the third intersection, which is this part, and this, people, is called as machine learning, right? Machine learning. Why machine learning? If you can combine the power of maths and stats, right? And you can combine the power of computer science, you will find yourself to be in a position where you can call yourself a machine learning engineer, right? How does a machine learning engineer become a data scientist? When they couple it up with the domain expertise, right? So in this course, people, in this course, we will teach you computer science, we will teach you a little bit about math and stats, but what we cannot teach you is domain knowledge. Yeah. And making the base of what we are about to do very, very important to understand, right? If you understand this, then half the battle is won, right? So now to answer the question which was posted earlier—answer to the question which was posted earlier—there are different things in the world of data science, right? So you can pick and choose anything, or you can do everything by yourself. If you guys are engineers, then I think you can be at the sweet spot going forward in life if you choose a domain for yourself. For example—example, you choose to be a part of the automobile industry, you choose to be a part of the medical industry, you choose to be a part of, say, the retail industry, you choose to be a part of the aeronautics industry, right? You choose to be a part of the finance industry, right? So whatever you will choose, this thing will get developed over time, right? This is the most difficult out of these three. I try to give you the example, right? My domain was the agricultural industry, right? The agri products—I have worked extensively in the agricultural industry, right? So again, for now, in my current role, this is something which I don't have—I have this expertise—I have this expertise—so same will be with you, and you guys will develop this knowledge over time. Now I will try to give you an example, right? Elections—to make you understand how data science can be used in one particular use case. Okay, maybe we can extend that to a lot of other use cases and examples, right? Talking about election season, right? Talking about the election season, right? We will—we will try to understand how do we use data science because it is very extensively used in this—data science, right? Now let me talk about the first phase. Let's say this is the pre-election phase, right? Right. This is the pre-election phase. In this pre-election phase, what do you think will be the tasks with which an agency like the Election Commission of India will be doing? The first task can be that they will be doing the voter registration, right? Voter registration and data management, isn't it? Yeah, it will start with that. And what will be the things inside this? The first thing will be data collection, right? First thing will be data collection. So you will collect the data from all the registered voters, right? Maybe it can be their demographic information: where they live, what is their age, what is their gender, right? What is their past polling behavior? Have they turned out previously or not, right? All these details we can collect. Then can we also do data cleaning because I've told you, right, that there can be a lot of redundancies. Some person's name can appear twice, right? Some people can be a mismatch. Suppose, for example, we have learned this in Python: Raghav. Raghav. Raghav. Raghav. Right. Right. All these are what, people? This is belonging to the same name, right? This is me. But for a computer, for a computer, how many Raghavs are there? Different ones. All are different, right? So examples like these, right? Some people might have died. They might not be existing anymore, right? So all this part will be taken care of where people—in the data cleaning process, right? Removing the duplicates, updating the new addresses, right? Correcting the information about every voter, all those things, right? And now lastly, we can also include a flavor of data analytics, right? What will data analytics include in this? We can analyze the demographic data to identify eligible voters, isn't it? Yeah. People—till now we haven't understood this—why it is not automated. How will you pick up all the—how—how will you pick up all the details and nuances? Suppose you're filling a form, right? By mistake you have—Suppose you are 25 years of age. Suppose you have written 250. So does that mean I remove this—I remove this entry of yours? Because by mistake you have written your age as 250. Is age 250 possible in our current world? Never, right? So I will have to tell the machine that because this is a mistake, please convert it to 25, isn't it? This is called as imputation. So this is mostly a manual task. Not manual, but you have to understand the problem manually and then code it on a computer. Suppose someone has written their state as E D L H I, right, and country as I D I N, right? So what is this state referring to? What is this country referring to? The humans are very smart—India, right? We can say this is India, and if this is India, then what is this pointing out to? Delhi. Just because you said—and it's a—it's a leading question—I wanted to explain this. Tell me—is it possible to do this automatically? No, right? We have to employ manual rules, right? We have to tell—because we—who is more intelligent—humans or machines? Humans, right? Machines are just more optimized, right? So we know through human intelligence that this is pointing to Delhi, and this is not EDLhi. So this is data cleaning, right? Lastly, we have data analytics. So do you think, people, based on these data points, we can understand that who are the eligible voters and maybe who are not registered yet, maybe who have not voted in the past? Can we do all those analytics and reach out to those peoples and persons? That's the first part. This is just the first part—pre-election phase. Now moving on to the second part, right? Moving on to the second part, let's say—we say—public opinion regarding the polling, right? The polling which is about to happen. The first thing will be—you want to collect information about people, right? So can—can you go out and reach all 1.88 billion people in this country that what is their likable vote for which party? Is it possible? 1.88 billion? Do you think it's possible? For 1 billion? Do you think it's possible? For 500 million? Do you think it's possible? For 100 million? No, right? So basically, I'm talking about—what?—I have something which is called as a population, right? And if I have to study about this population, which is 1.88 billion people, which is impossible—which you just said—what do I need to do? Should I stop my process? No, right? I will go and collect something which is called a sample, right? We always work in samples, right? Suppose someone says that a Coca-Cola bottle does not contain 500 ml of liquid which it claims, right? Suppose someone has put this allegation—possible that Coca-Cola bottles do not have 500 ml of liquid which they promise. Now there are two ways to deal with this, right? There are two ways to deal with this. Either I go and collect all the bottles of Coca-Cola in the world. Possible? Never, right? So what will I do? I will—

Go and pick up a handful of bottles. Right? Handful of bottles. So what is that handful of bottles? Those are called samples.

One last thing. Suppose someone says that because of an industry, all the fishes of the lake are dying or they are infected. Is it possible to go and collect and check all the fishes in the pond or a lake? No. Right. What will we do? We will collect again a handful of fishes and we will test them. Right? Again samples.

Now how does the raw data get collected? Raw data, as in I hope you understand this, this is no more about the population; with this step being told, you're now dealing with samples. So now, do you want to ask me how a sample is created? Yeah, now I'm coming to that. Now, guys, there are a lot of—the second point is how to sample, right? How to sample, right? So we have sampling techniques, people—one is called probabilistic and one is called nonprobabilistic. Right? I will not go into detail right now. I just want to tell you an overview.

Probabilistic is, suppose uh you are manufacturing t-shirts. Right? You're manufacturing t-shirts. Right? And suppose you created 100 lots of 1,000 t-shirts, right? So this is box one, box two, box three, box four, up to up to 100, right? 100 boxes. And in each box, how many t-shirts are there? 1,000. Right? Now suppose you are a quality inspector; you're a quality inspector. Is it possible for you to go over all 100 lots, checking all the 1,000 t-shirts one by one? No, right? What will you do? You will sample again. You will sample. Now the most common way of sampling in these kinds of problems is probabilistic sampling. What is probability? What is the probability of getting heads or tails when you flip a coin? Equally likely: 1/2 and 1/2. What is the probability of getting 1, 2, 3, 4, 5, 6 on a roll of a dice? 1/6. Now what is the probability of picking any t-shirt from this first slot out of 1,000 t-shirts? 1/1,000. Yes. So do you think all the t-shirts have equally probable—equal probability of being picked up without any bias? If you decide to draw five t-shirts, right, from each of these boxes, right, and suppose say two are defective and three are not defective. What will you do? Will you accept the lot or reject the lot? We have a majority of t-shirts that are non-defective, or let's say we will reject the lot. We will reject the lot. We will reject the lot. Let's say we will reject the lot. Okay. Though this is basically subjective as per the company policies, but let's say we rejected it. Now, people, my question to you is, what if this entire batch had only two defective t-shirts? But now what will happen? The entire batch will be rejected. Yes. Let's say you sampled one t-shirt. Let's let's change the use case. I say you sampled only one t-shirt and that t-shirt was defective. Now you will reject the batch, and that is an equally likely case. So this is called probabilistic sampling, people, and there is no way you can go back. There is no way you cannot say that, "Hey sir, please uh allow this batch to pass because there is a chance that the rest of the t-shirts are not defective." No, it is not the way that happens. It happens randomly. So, right, this is called random sampling. Right? Random sampling.

Now suppose you are doing cancer research, right? You're doing cancer research, right? So for your cancer research, people, what kind of people will you need? People who have had cancer in the past, isn't it? To know more about their problem, to know more about their medical condition. So is it possible, people, that in this use case you can go and pick up any person from the population and ask them questions? No. Right? That is not possible. So now is the probability equally likely, or has it changed when you pick the sample? It has changed. Now there is a bias which is introduced that you only want people who had cancer. Right? So that kind of sampling, people, is called nonprobabilistic sampling. Right? Nonprobabilistic sampling. Clear? Now sampling technique. Right?

Now the third thing in this same scheme can be, people, what? It can be the data collection, right? Data collection mode, right? That how do you collect the data? You can float a survey on, say, the internet. You can go and stand outside a mall or office, isn't it? How will you—how will you—can probably interview someone, right? Interview someone, right? You can have a group discussion. Yeah. All these techniques—yes, no, maybe, right?—in the same part, public opinion polling, right?

Now, guys, uh so this was a brief introduction, right? And this can be extended to any industry. Right? As of now, you can have an example in medical science, right? You can have an example in the automobile industry, right? You can have an example in retail, right? Right. Then you can have an example in manufacturing. Right? You can have an example in education, right? You can have an example in sports, right? IPL analysis, cricket analysis, all these sports analyses, right? These are the most famous domains, right? They are the most famous domains. They are not topics; they are domains, right? In which data science is used extensively. So, I'm just going to check. Guys, in automobiles, there's a biggest example: self-driving cars. Yeah, autonomous driving. How do you think that's possible? Data science again, like Tesla. Absolutely. Tesla is level three. We have five levels. Level three is narrow AI. Level four is AGI, and level five is super AI. We are going to move forward, right, with the technical aspect, right, with the technical aspect in our data science course, right, right, which is based on Python, right? Because that is our base language which we have learned so far, right? So in Python, people, we will start and cover four of the packages right now, and as we move on to machine learning and other uh deep learning and everything, you will explore more and more packages. The first package we have to cover will be NumPy. Right? I'll explain to you in detail what NumPy is. Then we will cover Pandas, then we will cover Matplotlib, and then finally we will cover something called SciPy. Right? We will cover something called SciPy. So these four packages inherently we have to cover in Python to make sure we are able to reduce the time in coding, right? We are able to reduce the time in coding, and using these packages immediately helps us in getting the desired results. Right? I hope you all remember the concept of modules we have covered in Python. Do we all remember functions and modules? You can use Jupyter Notebook. If your Jupyter Notebook is not installed, you can use something which is called Google Colab. Right? Go to Google, type Colab, right? Let's say you write Colab, right? And then you will see this option. Click on Google Colab, and it will allow you to code in Python, right? So we are now going to discuss about the NumPy package in Python. Okay, the NumPy package in Python.

So, NumPy is a fundamental package for data science in Python, right? It is one of the most fundamental packages for practicing data science in Python. Right? Why is that? Why is—so why is NumPy so fundamental? What's so special? The NumPy package gives us a new data type for handling data in Python called ND arrays. Right? ND arrays, which stands for—this stands for N-dimensional arrays. Right? N-dimensional arrays. Right? This stands for N-dimensional arrays. So till now, till now we have studied about lists, integers, tuples, strings, right? Out of which, out of which the data types such as lists and tuples have been used to store data, right? Store data, right? And range is used for generating new data which is primarily sequential. Right? Now there is one—now there is one problem, right? There is one problem, and there should be a question—there should be a question. Why do we need a new data type to work with data science, right? Why do we need this? Yep. So the answer to this question, people, the answer to this question uh lies in a small explanation, right? Yeah, lies in a small explanation, which is that Python is a high-level language, right? Python is a high-level language, right? And a high-level language is usually very distant from the hardware. A high-level language is close to hardware or distant from the hardware. Did you not attend the Python programming essentials? What is the type of programming language which is closest to hardware? A low-level language. If this is my hardware, yeah, this is my OS. This is my application layer, right? So high-level languages here and low-level languages here, right? Which is closest? So binary languages, assembly languages, right? All these are closest to the hardware because where is the processing happening? Where is the processing happening of the data? At the hardware, isn't it? Processing of data is happening at the hardware. No worry. Actually, it is happening at the hardware, right? What is processing? Processing is signals of zeros and ones, right? What are zeros and ones? These are electric signals. These are the electric signals, right? Which is basically on and off. Right? And it is communicated to the hardware through the help of resistors and microprocessors. Right? Isn't it right? Why a computer only knows zeros and ones? Because zero is off and one is on, which is the electric current, right? Electric current to activate or deactivate certain things, right? True and false gates, right? So, it is happening at the hardware. So, now, people, if you understand this part, then try to logically connect it to what I'm going to say. When you are studying data science, what kind of data will you be dealing with? Big data, right? And as the name suggests, it will have a lot of volume, right? It will have a lot of volume, other than a lot of other things, right? It will be very, very big. And on this large volume of data, you'll be doing processing. You'll be doing processing. What does processing mean? What does processing mean? Operations. So where is this operation happening? This is happening in the hardware. And for hardware, which is the closest language to hardware—a low-level language. But now, people, but now we have a situation in front of us. What is the situation—that what are we trying to do data science with? What are we trying to do data science with? Python, right? And Python is what? A high-level language, isn't it? Yeah. So, there's a discrepancy. Yeah. A big one because high-level—high-level language—these are slow in processing, right? These are very slow in processing, right? So for these kinds of languages to handle this kind of data and these kinds of operations, yeah, it's very difficult, right? So let's let's keep—let's keep this part aside if you understand this. Now let's go to the second point. Right? When you learned Python, on a scale of 1 to 10, how easy was it? The ease of use of Python—it's relatively a higher number, right? Relatively a higher number. So now, guys, when I talk about data science, right? When I talk about data science, okay? Right. When I talk about data science, my thing is that this will be used by the masses, right? Will be used by the masses. A lot of people—managers, programmers, business analysts, data analysts—possible, right? Who are from non-technical backgrounds, who don't know coding, they also can do data science because data science is a general thing, isn't it? Understanding the data. It should not be limited by your capability to understand the technicalities of a very complex language. So for these people, which language is suitable? Which is Python, right? It is easiest to understand. It's a high-level language, almost like English. Yeah. So Python is a simple language. So in this part, people, Python fits the bill, right? Which is a bigger thing. In the second part, when we talk about the operations—when you talk about the operations—in this part, there is a problem, right? Python fails, right? Python fails, right? Because it's a high-level language. We said that, okay, there is a language called C, right? Which is a middle-level language. Right. And C language, people, is used to create OS—operating systems. It is used to create networks, networking applications. It is used to create games, right? All the things which are close to hardware, C is used, right? So C fits this bill, right? C fits this bill. Python said that, okay, if C fits the bill and Python is written in C—written on C, right? It's written on top of the C language. Now what happened was Python said, okay, if my intrinsic data types—my intrinsic processing—is not suitable for data science, but my high-level nature is—let's do one thing—let's take the C language and use its power, right? Use its power—that it is very close to the hardware—and let's create a new data type, right? Let's create a new data type which is written on top of C and which can integrate with Python seamlessly, and people, this new data type was called an array, and this was given to you by something called the NumPy package, right? The NumPy package. It defined a new data type which was array, and along with defining the array—array—it gave various operations, right? It gave various operations one could perform on arrays, right? One could perform on arrays. It's simple, right? The—the power of processing lied with C, right? It was lying with C. So we developed a new data type using C on top of Python, and that new data type was called an array. And this array was defined in a new module which was called the NumPy module, which told you how to create the arrays and then how to manipulate those arrays for doing data science. We had options like Java, we had options like C, right? We had options like Fortran to be used for data science, but we chose Python because of its simplicity and the simple syntaxes that people from non-programming backgrounds could also use Python to do data science. Right now the limitation was that because it is slow because of being high-level, we needed something which could make it fast, and that was using an external data type which is not internal to Python, and that was array, and this array is defined inside a new uh module or a library called NumPy, which is numerical Python, right? Numerical Python. Back to the programming, right? Where we have understood, right, about this question, right? So NumPy essentially is built on top of the C language, which is compatible with Python. It leverages the power of closeness of C with hardware, right? C with hardware, which eventually makes the processing faster in Python. Right? So this is what the first part is, right? This is what the first part is, right? Now the second thing is, people, so this is about the performance bit, right? These are the performance bits. The above is about the performance of—right? Now coming to the next part, which is the memory efficiency, why, right? Memory efficiency, right? So arrays created by NumPy in Python are less memory exhaustive than lists in Python, right? And I will prove these points later on to you through code, right? In lists, right? In lists, each item is an object, right? Each item is an object, right? I hope you remember this, guys. Each item is an object, right? And it holds, right? It holds meta-information like references and types, right? Etc., right? A lot of information it holds, right? This makes lists consume more memory, right? But but arrays in NumPy are contiguous, which means that they do not create objects but rather directly store the data in continuous memory blocks, one after another. Right? Also the arrays are homogeneous in nature. Right? You can only store homogeneous data in an array, unlike lists. In lists, you could store different, different data types. Right? But in arrays you cannot, right? You have to store the same kind of data in the array—homogeneous. So it's a contiguous memory block, which is meaning that you can store data in continuity, right? One after another in the memory block. So the access is faster, the memory location, and the memory efficiency is very higher, right? As compared to the native data types like lists or tuples in Python. Arrays are basically vectorized operations, right? I'll talk to—talk about—to you with vectors—what are vectors, right? They are the vectorized operations, and they are way more convenient to deal with as compared to lists in Python. Right? In lists, you have to go through a lot of loops, right? We saw that we have to go through list comprehension, but you will see in Python Pandas—sorry, in Python NumPy—the vectorized operations are very, very simple, right? They are very, very simple, right? Again, for all this, I will give you examples, but uh it will take some time because you have to understand what arrays are first, right? Okay.

Then the most important other packages which are Pandas, Matplotlib, SciPy, Scikit-learn, all are written on top of the NumPy package. That's the reason—for this reason—it's called a fundamental package. All the other packages which make your life easier as a data scientist, where you don't have to worry about the code—all these packages make our data science journey smooth because we have to worry less about code and more about logic. Yes. So for these—for the understanding of these packages—it's very important that we understand NumPy first and then we move forward. Right now, then let's get started. So the first step, right? The first step which you have to uh see, right? The first step which you have to see while using the NumPy package is basically from where will you import the package, right? From where will you import the package? So to import the NumPy package, we write `import numpy as np`, right? Where `np` is an alias, right? It's an alias, right? So please do this: `import numpy as np`. Right? If you don't get any result for this, then you can write `pip install numpy`, right? `pip install numpy` and just execute this, right? When you will execute this, it will give you this message or it will give you—it will download this package for you, right? Pip stands for Python Index Package, right? It is basically like the Play Store, App Store, right? Or the Windows Store for Python, right? So you're just going to these Play Store, Windows Store, App Store of your Python and asking them to download this for you, right? It is also called the package manager, right? Pip. If it is done, you can also check the version; you can say `np.version`, and it will give you the version of NumPy. Right? NumPy is an open-source package, right? Yeah, that's about it. Yeah, it's an open-source package, people, right? Open-source package. And if you want to see the code, you can go to GitHub. Go to Google, write the uh code for NumPy. It will show you it on GitHub. Right. Done. So let's say I—okay, so I say we have—when we—so okay, before this, let me come to a little bit of theory before I do this with you. So now, guys, I said that NumPy has arrays as a data type, right? And this is written on C, which runs directly on hardware. Also, this NumPy array is basically already a vector, right? It's basically a vector. Now, what is a vector, people? What is a vector? A vector is a quantity which has a sign and magnitude, right? It has a sign and magnitude. If I say this is a Cartesian space, this is I, this is J. And I say this—this is 3I and 4J, right? 3icap, 4jcap. So this is a vector, right? And this is the direction, guys. If you have studied elementary math, you would know this. Yes, people, this is a vector. If I draw another like this, then this is another vector. So I will call this, say, uh 2I and 5J. Yeah, this is another vector, and this is theta, right? This is the direction. This is the direction. And what is the magnitude? 3I—3² + 4², which is 9 + 16, which is 25, under root, which is 5. So the magnitude of the vector is 5, and the direction is equal to theta. This is a vector quantity. Now what is this? What is this? This is a scalar. This is a scalar. Yeah. Only magnitude, isn't it? This is a scalar—only magnitude. And then I have 2, 3. Right? This is what? This is a vector. It has two dimensions or one dimension. Only one dimension, right? This is a one-dimensional vector. Yep. One-dimensional vector. Now if I say this: 2, 3, 4, 5, what is this called? This is called a matrix, right? This is called a matrix, which is what? A collection of vectors, and this collection of m vectors—matrix—is called two-dimensional, right? It is called two-dimensional. Now if you have this—so these are stacked behind each other. This is one. This is two. This is three, right? So we have three layers in the matrix. So how many dimensions will this be, people? One dimension, two dimensions, and three dimensions. This is a three-dimensional matrix, right? Or a three-dimensional vector or a three-dimensional array. I'll repeat once again. What is a single value? A single value is called a scalar, right? It only has magnitude. Now when you...

Have multiple values, right? This is called as a vector. It has a direction; it has a magnitude. And this is single dimension. Now multiple vectors, right? Like this, or maybe you can say like this, right? Are you understanding why I'm calling it one dimension? It can be either this dimension or it can be this dimension. In any dimension you stack two vectors, you will get yourself a matrix. Right? You will get yourself a matrix which is now two dimensions. It has rows and it has columns. Right?

Now if you stack multiple such matrices one after another, right? This becomes a three-dimensional matrix. And this can go up to how many dimensions, people? How many dimensions this can go up to? It can go up to—this can go up to n dimensions, right? N dimensions, right? So now do we get it? Why do we call it ND arrays? Yeah, n-dimension arrays, right? That's why are we calling something what? Right. N-dimension arrays. We are only capable of viewing three dimensions, people. It can go up to 100 dimensions, 500 dimensions, 1,000 dimensions, any dimensions. Scalars are least important; vectors are more important.

So now we have multiple dimensions in arrays, namely 0D, 1D, 2D, 3D, and so on till the ND, right? ND arrays. Right now let's create our first array. Right? Let's create our first array. And this will be a zero-dimension array, right? Zero-dimension array. How will you create this? You will say arr0 is equal to np.array and you will mention what a scalar value is. What is a scalar value? It is a simple value. I say two, right? np.array equal to two. And when you will now check or print the type of arr0, it will tell you class numpy.ndarray, right? It is a zero-dimension array. And if you want to check the dimension, you just have to write arr0.ndim, right? And it shows you that there is zero dimensions present. So what is this in short? This is a scalar. I have entered a single value, people—2, 200, 500, whatever you want to enter. And the syntax is np.array. np.array. You're instructing numpy package to fetch the function array on method array and convert this into that particular data type, right? array zero, and when you check the type, type is ndarray, but what is the dimension of this ndarray? This is zero, which is nothing but a scalar, right? We have created this; this is what we have created.

Right now, guys, if I created this list, right? Say lis is equal to… Yeah. So this was—this was a list, right? This was a list, and this was—this is what, people? What is this that we have just studied, according to that? What dimension is this list? So what I'm trying to do is I will create a one-dimensional array with a—or let's say from a list, right? Let's—let me show you how do we do that. Okay, so I say… Hurry, read the error: NameError: name 'arr' is not defined. Why? Because you have created array from a arr0. Come on, hurry. Right. I will create a one-dimensional array. How will I do that, people? I will say arr1 is equal to np.array. And can I pass a 1D list inside this? Or can I say I can pass lis inside this? When I do this, people, now see what will happen. arr1 will be equal to this, right? And if you say—let me say print arr1, it will be like this. Okay, this is your arr1.ndim; you will see that it gives you one, right? This is a one-dimension array, right? One-dimension array. Yes.

So, for example, when I say scalar, right? When I say scalar, I say 35s, right? When I say vector 1D, I say… uh… so these are suppose my marks, right? Now I say 35, 40, 50, right? So now what are these? My marks in three subjects. Yeah, marks in three subjects. This is my say Hindi, this is English, and this is maths, right? Or let's say science, because not everyone has Hindi, science, English, and maths, right, guys? Now if I have to create a matrix of two dimensions, what will I write? So that means this is one student, this is one student, people, isn't it, guys? Yes, no, maybe? So now in a matrix, people, we will have what? We will have multiple students. Yes, no. Maybe. In a matrix, people, we will have multiple students. Suppose this was S1. Now you will have S_1, S_2, S3, S4 like this. And each student will have their own individual list of marks. So can I say that I'm making a nested list? Can I say that, people? I'm making a nested list. So now let's make it okay from a nested list. Okay, a nested list. So I'll say arr2 is equal to np.array(list). I will have to create a list first. lis2 is equal to… So this is my first bracket. What is this bracket representing? This bigger bracket. Now I will put another bracket inside this and I will write 1, 1, 2, 2, 3, 3, 3. I'll put a comma again. Write a comma. Then I will say 4, 4, 5, 5, 6, 6, comma, 7, 7, 8, 8, 9, 9. Right? I'll do this right now. I will say list2. When you will do this, you will see that an array like this has been created, right? Like this. arr2, right? This has been created, right? When we check the dimension, it is a two-dimension array, right? Yeah, marks of three different students in three different subjects. And this is the same technique; you can create a three-dimension array. How will you create a three-dimension array, people? If I go here, how will you create a three-dimension array?

Now suppose I have data in this, and this is my master list. Okay, this is my master list. In this, I have data, and this can be represented like this. This is my first matrix, isn't it? And this is the vector inside this: V_1, V_2, V3. Then this can be called as M1. And now to create a three-dimension setup, how many M1s do you need? You need multiple M1s, isn't it? You need M1, M2, M3, multiple matrices like this, people. Like this: matrix 1, matrix 2, matrix 3. So what will you do? You will have—you will have what, people? You will have another yellow, right? And you will have inside this yellow multiple purples. Isn't it, right? Again like this. Yeah. Like this you will have it, people. So can I say, people? Can I say that as I am increasing the dimensions, as I am increasing the dimensions, I am putting 1D—sorry, 0D, right? And… okay, again in this V1, in this V_1, do you think you will have multiple scalars, people? Can I say that? Can I say multiple scalars create a vector, multiple vectors create a matrix, and multiple matrices create one three-dimensional matrix? Can I say that?

Let me talk to you about an image. Right? Image, right? What is an image made up of, people? What is an image made up of? Of pixels. Yes or no? No, not frames. Frames are basically videos. Pixels are creating an image, right? So we have how many pixels here? 1, 2, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16. Right? Suppose this is one vector, right? This is one vector, and this is one scalar, right? Scalar 1, scalar 2, scalar 3, scalar 4. And this will make vector V_1. This is V_2, V_3, V_4. And together—together can I call this M_1 and call this blue? So for a colored image, how many channels are there, people? How many channels are there? What do we call it? We call it the image as RGB—RGB image, that means red, green, blue. So this is the blue part. So similarly, you will have a red part in front of it. Then you will have a green part, and then finally you will have a blue part. Yeah. Are you understanding, guys? Why do we require three-dimensional arrays? Yes. Suppose now you want to make a change at this—this pixel. So you will go to the third layer, which is the blue layer. Then you will go to the third column. You will go to the third column and third row. And this is how you will reach this pixel. Everyone? Yes. No. Maybe. Yes. So this is the reason why we need to create a 3D array. Right? To ingest information like this. Right? To ingest information like this. So we can create a 3D array also. Right. Right. And how did I tell you? How many brackets will I have? Squared brackets. I will have three squared brackets. So now this is just one student. Right. Now I'll put a comma here. Right? I'll put a comma here, and I will start. Right? I'll do this. I'll say this list three… arr3… arr3… r3… r… and now you will see that there's a three-dimensional array. Right, guys? Right. This is a three-dimensional array. Yep.

Moving on. There are multiple ways to create arrays. Right. The first one we have done. So we have done from lists. From lists we have done. Then second will be from—uh—we can create a zero array. Right? We can create ones array. Right? Then we can create custom array. Right? I will show this all to you. Right? I'll show this all to you. So let's start with the ones array—sorry, zeros array, right? What do you have to do? You have to write… So let's say zero_arr0… okay, a zero-dimension zero array. So you say np.zeros. Right. And you create a two, right? A two, right? And if I say uh… zero.ndim, right? You will see it is a one-dimension array, right? It is a one-dimension array, right? So two is by default taken as… let me just show this to you. It will look like this, right? It is taken as horizontal. What is the dimension of this, guys? A vector. What is the dimension of this vector? No, no, no, it's one. It is basically 2 + 1, right? 2, 1. Sorry, 1, 2. My bad. 1, 2, isn't it? Now what if you had to create a 2 + 1? What if—what if you had to create a 2 + 1, right? So let me just show that to you. So I say… wait… like this. Okay. Now when I do this, I say arr_ones and I say here 2, 1, right? I say 2, 1. Now you will see, people, what will happen to this now because you have created—right?—specified two dimensions, right? What will this be converted to now? Hm. What will this be converted to? This will be converted to a two-dimension vector. By default, it was this, right? Which was this vector, right? By default, it was this vector, right? What is the dimension of this vector? 1 + how many values you put here, right? 1 + 2 like this. So we call this only one dimension. We call this only one dimension. But now when I will run this one, you will see… yes, 2 + 1. And now you will see this is two dimensions. Right? You see this is two dimensions now. Clear, people? Just the orientation has changed. But now you see the brackets; there are two brackets now because what have you now instructed? You have now instructed Python and rather numpy to create the vector as 2 + 1. So you have said, give me this zero and give me this zero here. So the moment you do this, you are now specifying the rows and columns. Now this 2, 1, right? Let's say this is 2, 1. Now let me create another 2 x 2 dimension matrix for you. Let me call this 10, 10. Right? So how many rows and columns will it have, people? How many rows and columns will it have? I will say 2, 1. It will have 10 rows and 10 columns like this. You saw this? Yeah. 10 rows and 10 columns. Right? By default, it is 1 + 2 like this. It is a one-dimensional vector, and this is a two-dimensional vector. What about a three-dimensional vector? zero_… arr3 is equal to np.zeros, right? np.zero… uh… should I just write 3, 3, 3? And I should check for this. Yeah, you will have a 3 x 3 vector—3 x 3 matrix with three matrices stacked behind each other. Right? If I say four, this will be four. So how do we read this? Number of layers, number of rows, and number of columns. So can you help me with a syntax? Can you help me with a syntax which can create me a 3D matrix of five layers, three rows, and three columns? What will I write? Five layers, three rows, and three columns. What will I write? 5, 3, 3. Yeah, you'll get five layers. 1, 2, 3, 4, 5, right? Like this. Suppose—suppose you have to represent an image with—with say RGB channel and uh… say 256 x 256 pixels. How will you create this? So I'll say img—right?—image is equal to np.zeros, and I will say inside this three channels, 256 x 256, right? And when you will just run this image, it will be like this, right? It will be like this. This is one channel, this is two channels, and this is three channels, and uh… we understood that numpy is a fundamental package for data science in Python. Right? This package gives us a new data type to work with, which is called as n-dimensional arrays. Right? Now what are arrays? What are arrays? Arrays are nothing but vectors, right? Which are stored in a contiguous block of memory, which means they are stored continuously one after another, and they do not get converted into the object unlike the list, and they are way faster; they are more memory efficient than lists. I will prove this fact to you today with the help of an example, through the help of code, right? So why numpy arrays? Because numpy is a package which is built on top of the C language, which is compatible with Python, and C being a middle-level language interacts directly with the hardware. So whatever operation you run—in fact, that is getting directly executed on the hardware itself. Right? That is the reason why the uh… performance is way better when we try to use the n-dimensional arrays. Along with this, these syntaxes—the type of syntaxes we used to type uh… in lists—I will show that to you also today with the help of an example—are way simpler when you try to do them with ndarrays, right? So arrays are basically vectorized operations, right? And they're also very convenient to deal with as compared to lists and other native data types in Python, right? And the most important part is that in our data science journey, whatever other packages we will use. Right? Again, what are packages? They have predefined things stored for you so that you can leverage them and focus less on code and more on logic. Right? You need to be aware about the logic more than the knowledge of the code. Right? So we have packages for that which contains methods inside them which you can use and uh… without any further calculations you can work with them directly. Right? So that is how we started with numpy, right? And the syntax to import numpy was import numpy as np, where np was an alias, right? Uh… you could do pip install numpy if someone did not have access to numpy—if numpy was not coming by default. You can use pip install numpy, which is Python Index Package, right? And uh… this is like Play Store, App Store, Windows for Python. All the packages are stored in pip, and you can call pip—you can ask pip to download that package for you so that you can use that. Right? It's basically the package manager. So numpy is an open source, and if you want to see the code, you can go to GitHub and check the code out for yourself, right? In numpy, we have n-dimensional arrays. Now the question arises, people, that why do we need numpy, right? So my answer to this particular question is that when you will deal in data science, right? When you will become a data scientist, you will be dealing with data. Right? Now my question is, how will you ingest—how will you make the machine ingest the data? Right? There has to be a way, right? For you to input the data to the machine, right? To make manipulations to the data, to read the data. So all this is started as the base package of numpy arrays, right? Other than that, it becomes very difficult and cumbersome for us to deal with that, and this provides us a lot of ease and flexibility to deal with such massive amounts of data, which you will see along with me as we will move forward in this particular course. Yeah, perfect.

Now, people, let me just uh… pull up the PBTs. This is what we're discussing, people. We start with something which is called as a scalar, right? We start with something which is called as a scalar, which is a quantity which only has magnitude. Right? In numpy language, this is also called as 0D, right? It is called as zero dimension. Then, people, we have vector, right? And vector has a constant dimension. It has only one dimension, right? You can interpret it as a row or you can interpret it as a column; it doesn't really matter because this is only one single dimension, right? So usually it will be written as 5, blank, right? There will be nothing written in front of it. So this, in numpy terminology and nomenclature, is called as one dimension, right? When you move on, then you get—combine couple of vectors—you get a shape, and now that is called as a matrix, and in numpy terminology it is called as two-dimension, right? And when you try to stack multiple two-dimension matrices one before—with one after each other or one before each other—then they become something called as three-dimensional, and now in my capability I don't know what four dimensions look like, but there is a high possibility that you have n-dimensional data, right? It has—it has n—we are dealing with n-dimension data, right? Suppose with this I also add time, right? At t = 1, at t = 2. So that will serve as the fourth dimension for this data. But how do—how does it look like? I don't really know that. Right? Because humans are only capable of visualizing 3D—three dimensions at max, right? So you can go to n dimensions, and hence the name n-dimensional arrays. Right? ND arrays. Post this. Right? Post this, we moved on to create certain things, and I tried to explain you the data. Right? So the data will look to you like this. Right? You might have a scalar quantity which is marks, right? One mark, right? Now if I go on to vectors in one day, it can be marks of one student, right? Marks of one student in science, English, and maths, right? 35, 40, and 50 out of say 50, right? Three subjects. So this will be characterized as a vector. Right? What will be the dimension written? For this, it will be 3, nothing. This will be zero, right? Shape will be zero. For this matrix, suppose we have 1, 2, 3, 4 students, and each student will have three marks. Right? Each student will have three marks. So what will be the shape of this? We have four rows and three columns. Right? So this will be the shape, right? Of this 2D matrix, right? And now if you stack images, right? One behind each other, then it will be like this, right? Image—I gave you an example—so this has 4 x 4 pixels, so the shape will be 3 x 4 x 4, right? This will be the shape for this particular 3D matrix, right, guys? So this is how you input the data. Just to tell you a little bit more, since generative AI is very popular these days. Right? So what if I tell you the fact that the ChatGPT which you use or you might have used, right? Has never seen a single word in its lifetime. Right? All it sees is numbers, right? Only numbers. How do we see numbers? Suppose I say my name is Raghav. Raghav is a nice name. Right? So these are two data, right? These are two data points. Now we all know that computers do not understand these, right? Computers do not understand these. Right? There is nothing—no understanding for computers to know what text is, right? It only knows 0 and 1. Yes, Dhika, right? It only knows zeros and ones. So see how we will convert this. So there is something called as vocabulary, right? So vocabularies are nothing but the unique words, right? How many unique words do I have in this: My name is Raghav. Four. This is not unique. This is repeating. This is repeating. Five, six. And this is repeating. So I have six words. So now, guys, I will do something called as word-to-vec, where I will convert these words into vectors. How will I…

convert them? Look at this. So I will have, suppose this is S_sub_1, this is S_sub_1 and this is S_UB_2, right? So I will represent S_sub_1 as, right. I will have a vector of size six. How? I will say 1 0 0 0. Right? 1 0 0. How many elements does it have? Six elements. Right? Name will be 0 1 0 0 0. is will be 0 0 1 0 0 0 and ra will be 0 0 0 1 0 0. Right, this is s1. My name is Rahab. Now when it comes to s_ub_2, right, when it comes to s_ub_2, how will I enter this s2? 0 1 0 0. Ragh, what is is here? 0 0 1 0 0. What is 0 0 0 1 0? Or nice is 0 0 0 1 and name will be 0 1 0 0 0. Right. Now this will be the vector representation of these two sentences. Just to tell you a fact, GBD3, right, GPD3 model, right, GPD3 or GPD 3.5, they have vocabul, vabulary of 30,000 words, right? 30,000 words. And each word, right? Each word has a dimension of 12,500 numbers. Right? What do I mean? Suppose I say Raghub. So it will be one word and it will be represented by 11,500 different numbers, right, and like raghub there will be 30,000 words in this GPD model, right, 30,000 words in this GPD model, right, and this total number of parameters which get trains in the neural network, which we will learn later on neural networks, they are almost close to 1.75 billion parameters, right, 1.75 billion parameters. So why I'm telling you all this, because to showcase to you that what is the importance of vectors, right, in the entire machine learning and data science, yeah people. So this is the reason people. Now my question is, how will you create these vectors? How will you read these vectors? Right? The answer is through numpy package because it is the base package. Clear people? Yeah. I hope today's class will be uh interesting for you because you will know the context. Why are we doing it? Yeah. So I'll try to show that to you how we convert things to vectors, right? Okay. Let me go here now. Right. Let me go here. So people, uh we started using numpy. So I started with the zero dimension arrays, right? Zero dimension arrays. So zero dimension is nothing but a scalar. So I created a ar r0 which was np.dot array and I entered a single word, single element inside this which is nothing but a scalar. And then I checked the type of uh ar r0 also, right, and then I check the dimension also. So the answer was, type was numpy nd array and the dimension was zero. Right? Exactly what I had mentioned in my DPD, right, it will be having a zero dimension like this, right, same thing has been proven, right, because it's a scalar. Now coming on to one dimension, right, coming on to one dimension, I create a list which is nothing but a one-dimension data type, right. Now I create an array which is a ar r1 from array from this particular list lis, and then I check the type of print ar1. Check the type of ar1 and the dimension. Right? So when I execute, right, then you will see that it was this numpy array and the dimension was one, right, exactly like this. So if I show you something else, say print a ar r1.shape, you will see it's 4, a empty, right, and I show it to you here, it will be empty, right, empty and this will be comma one, right, so this means that it has only four elements. Right? If I increase these elements to say 55 66 77, then it will become 7, blank, right? Which means it has seven elements as a vector. Now we create something with a nested list, right, which is like this. So with a nest, I want one bracket which is running outside, right, then inside this I have one, two and three lists inside one list. Right? So this is a nested list. This is marks of first student, second student and the third student. Right? So I do this and you see it is this, right? And I will show you the shape also, right? It will be 3, 3 rows and three columns. Three rows and three columns. Right? Now similarly we can also create a the 3D matrix, right, with two levels, right, level one and level two and 3 + 3, so the shape will be what people, can someone guess the shape, what will be the shape of this I've shown you here? Right? If this is 3, 4, then what will be this? Three rows and three columns. Right? So when you will execute you will get 2 3 3. Right? 2 3 3. Right? Right now people, there are multiple ways, right, there are multiple ways to create arrays and we should know them because all of these comes very very handy. Not right now. I don't have enough context to give you right now. But later on you will see with me or with some other trainer that how these will be used in deep learning specifically, right? They are the key of deep learning algorithms, right? Where we initialize some weights, we initialize some biases and those initializations are nothing but multi-dimensional numpy arrays, right? Numpy arrays. Okay. Like for example, suppose I have this, I have to multiply this with some random numbers, right? So how will you generate these random numbers? We will generate them through numpy. And you can generate them in a specific kind of uh shape, right? Which is 2 + 3. And then you can multiply them. You can multiply the matrices and you can get your output for yourself. Right? So this is the way they are used. So we saw the first thing from list we have already covered. Then now we are moving on to creating zero arrays, right. So I create a zero array of one dimension, right, of one dimension which is 0 0. They are represented in floats, right. They are represented in floats 0.0 zero. Right. Now you can create a two-dimensional zero array. Right? You can create a two-dimensional zero array which is you have to mention just the shape inside 2 + 1. So it will have two rows and one columns. Right? Two rows and one columns. The difference here is, the difference here is that these are one dimension and these are two dimensions. Right? You have explicitly mentioned the rows and columns. So you can expand this to 10 + 10 also. Right? You can expand in 10 + 10 or 10 + 6 whatever you feel like yourself. Right? It will have 10 rows and six columns. Right? Now you can also create the 3D arrays, 3D zero arrays. Right? Which is 5a 3a 3. What does five means? What does five means? First element represents the number of layers. So you have five layers, right? It's five layer deep. Then you have three rows and three columns, right? So it will look something like this. 1 2 1 2 3 4 5, right, like this something like this, right, it will look like this tab 1 2 3 1 2 3 1 2 3, right, 5 3 3. Okay, yes, the number of matrices, hurry, what I represented, right, layers rows columns, right, layer rows and columns. Clear? So now when you execute this, you will get an arrangement like this. Okay. I tried to show you this thing with another example, right? Which is I created an image of an RGB image of 256 cross 256 pixels, right? Which have all zeros inside them. And this is how it was created, right? Three layers RGB 256 256. So this is how the image will look like, right. This is how it will look like. Now you can also create arrays with ones. Right? Exactly the same way you created it with zeros. I'll give it to you. I'll give you 5 minutes time to create them. I'll show you one. So I say a1 is equal to np.ones and inside I pass two, I check A1. So this is a array like this, right? It is an array like this. Now you create, create two dimension and three dimension arrays of one. It is basically as a float. Hurry. It's represented as a float. Right. It's represented as a float. Okay. Right. Perfect. Right. Also guys, uh with this, right? Also with this you can create the custom arrays. Right. You can create the custom arrays. Right. How do we create custom arrays people? How do we create the custom arrays? You have created now zeros. You have created now ones. Now what is left that you create the custom arrays. Uh forget about this. I will come to this later on. Right? Let's create custom arrays. Right? So the syntax remains the same. Right? I'll say cus arr is equal to np. Right? This is the syntax np. Right? And you will say 6 + 6 and suppose you want an array of all fours. Right? This is the dimension 6 + 6. And this value after comma is basically the value which you want. You execute this and you copy this, paste this and you will get the arrays of fours for yourself. Right? If you want of 10, you will get of 10. If you want 10.3, you will get 10.3. Right? anything which you want. If you want case, you will get case, right? All the examples. So, let me just show that to you. Yep. Like this. Now guys, how did we create or how did we use range in Python? Can you use range to generate numbers between 20 to 50, right? 20 to 50. Can you give me the syntax quickly? How did you do that in range? How we do that? We said r is equal to range 20 to 51. Right? And then I said i in r print i. Right? And this is how I got the numbers. Right? So similar to range in Python, we have a range in numpy. Right? How do we use a range? I say uh a range ar is equal to np.a range, right, and then same syntax I will say 20 to 51, right, 20 to 51 and that's it and when I will check my ar range error you will see I have generated myself numbers between 20 to 50 and a range in numpy, right, a range in numpy. Yes, if you want a interval so you can use this, say a range one and after comma you pass the third argument. Suppose it's three. So now it will jump three times, right, 20 23 26 29 32 35 like this up till 50. Now guys, there is something which is called as lin space. Right? What is lindspace? It stands for linear spacing, which means between two given numbers. This function will fit the required number of numbers. Right? For example, suppose for example, we need to create an interval from 0 to 1. People, there are infinite numbers I can have between 0 to 1. Isn't it? Infinite numbers I can have between 0 to 1. 0.000000001 0 0 0 1 0 0 0 1 0 1 0 1, right, I can go in the infinite manner, right. Now for example, you need to create numbers between 0 to 10, right, and you want to create and want to have 10 numbers in it, right, so how will you do this? It's not float, it's about the number theory, right, between 0 and one you have infinite numbers. Right? So you say lspace is equal to np.lindspace. Right? np.t lindspace. You mention from 0 to 10, you want to have 10 numbers. Right? And when you will create lindspace you will see that these are the numbers are there which have been created. Right? These are the numbers which have been created. Right? Nine numbers. Now I say 100 numbers, I want evenly spaced 100 numbers, right, evenly spaced 100 numbers. How are they even? You can simply subtract one number from another and the difference for all the numbers will be exactly the same, 0.1. Subtract any two numbers, it will be 0 1 0 1 0 1, right? Where do we need this? We need this to plot the axises, right, when you will plot graphs you will need access between this interval, you need five values that works like this. Okay. Suppose you want from 0 to 10, five different values, right? You will have five different values like this, right? Between the gap of 0.5, right? If I say 1 to 10, you will have values like this, right? Like this. Clear? 100 values like this. Yeah. Clear guys. How do we use lin space? Suppose you want from 0 to 10. Interval from 0 to 10 and 10 will be included. 0 will not be included, right? We'll start from one. Why I go to one? It will be from one. Why is 0 not included then? Yeah, 0 is included, right? 0 is also included and 10 is also included. 100 numbers between them, right? Clear? This is what lind space is. Now guys, now suppose you location l space is basically used if you want to create n numbers between the range of numbers, right, between 0 to 1, right, suppose between 0 to 1 you are trying to plot a graph, okay, and your values are 0.2 0.3 0.6 six, right, and you want to draw a graph. So you will have to mark the axis, right, the x axis and the y axis, so you can use lindspace there and what will it do? It will take the range, it will take the interval in between you want to add the equal space numbers and then the third argument here will be that how many numbers do you want between them. So this syntax tells you that from 0 to 1 give me 100 numbers. How are these numbers? Equally spaced numbers. Equally spaced numbers. Right? So when you will execute this you will see that all there are 100 numbers which have been generated. Right? 100 numbers. And all the numbers are equidistant from each other because difference of every single number from the next number is 0.01 0 1 0 1. Yep. That's what it does. Right. Now guys, now suppose we want to generate random numbers, right? We want to generate random numbers. Right. Now we interested in generating random numbers. So we have something called as random.random. Right? What will it do? Random.random will generate random float numbers between 0 to 1. Random float numbers between 0 to 1. How will this happen? You will say rand rand is equal to np.random.random and inside you will mention what is the dimension that you seek. Suppose I want 6 + 6. So when you will check this you will have all numbers for 6 + 6 dimension, right, 6 + 6 matrix, right. Now every time you rerun this the numbers will change because all of these are random numbers, right, all of these are random numbers. Now I say 100 multiplied by rand rand, you will see all of them, all these numbers will be multiplied by 100, right, all of them in one shot. Now just like this we can also create random integers, right, how will we create random integers guys? I say rand intore rand is equal to np.random.rand, right, and here you will specify that what is the range of numbers you want from, so I say between 20 to 25 I need random numbers and then I want it from in a 3 + 3 format. Right? And now when you will check your random you will get random numbers generated like this. Okay? Random numbers generated like this. Right? If you say 3 + 3 + 3 you will get a 3 + 3 + 3 matrix. Even if you will only say 3, you will get a 1D. Right guys? You can change the dimension. So this is the range from which you want to choose the random numbers and this is the dimension you want this matrix or vector to be in. Now guys, we'll move on to the next part which is basically properties and again there are a lot of operations you'll have to see it yourself, right, properties and attributes of num py arrays, right, property and attributes of num py arrays. Okay. Now guys, the first one in this scheme of things is shape of array, right? Shape of array. What is shape of array? It tells you the dimensions of the array stored in a couple. Right? For example, I say a ar r0, right? A R R 1, A R R 2, A R R 3, right? And then I say print this.shape, right, like this and you will see that it will give you the shape of each array, right, 0D 1D 2D and 3D, right, people, shape of the array, please try it out, we have used it one or two times but this is what shape of array actually means? Second is people, nd, right, is nd, right, nd of array, right, it tells you the rank of the array, whether it's one dimensional, two dimensional, zero dimensional, three dimensional, four dimensional, so again I will do the same. Okay. Right. I'll say nd and you will see it will give you 0 1 2 3, zero dimension, zero rank, one rank, two rank and three rank and it can go all the way up to end rank. Before this, I should have also printed these arrays, right? These are the arrays. Yep. These are the arrays which we have and these are the subsequent things, right, copy array. Then guys, the third thing is the size of array, right, it tells you the number of elements inside. How many elements do we have inside this array in a ar r1? How many elements do we have? 1 2 3 4 5 6 7. How many elements do we have in this 2 + 2 A R2? 1 2 3 4 5 6 7 8 9, which is rows multiplied by columns 3 * 3, right, and how many elements do we have in this 3D which is 2 * 3 * 3 which is 18, right, so now when you copy this, right, you can use this and say size, right, it will say 7 9 18 be. Yep. Fourth is people. The D type of array tells you the data type of the array. Right? And I've told you we place only homogeneous data in array. Right? What will happen if we don't do this? I will show that to you also. Right? So we do this. Right? And we say retype and you will see in 64 all of them are integers, right? All of them are integers. That's the reason we are getting in 64, right, suppose I create a new array a ar r rore new, right, let me say head, right, hetro, hetrogenous and I say it is like np.array, let's say like this, right, like this. Now people, when you will check ar r.d type you will see it will give you float just because of one floating point number inside this entire array it gets converted to float, right, between all the integers if you put one float then it will be taking float directly, right. Now let me just copy this and let's say heterogenous one and let me add another value which is string and say rather, right, and when you will execute this it will give you U32, U32 here is representing strings, right, it is all called as objects, right, these are all string values, right, so precedences strings greatest then float and then your uh integers, right, if you place the heterogeneous data inside the numpy array, right, you only need to put homogeneous data in the array. Now fifth is the item size of array. Right? What is item size? It gives you the byte occupied by each element of an array. Right? Because we assume that elements will be homogeneous. It will give you the byte occupied by each element of the array. Only one element. Okay. So how will it happen? So let's say a ar r r0 or they say ar r r1.item size, right, and you will get eight, right. So why eight? Because because each data point, right, each data point is occupying the result is 8 bytes. Let me put this here. Right. Eight bytes because each element in ARR1 is occupying 64 bits which are equivalent to which are equivalent to 8 bytes. Right? One bite is equal to 8 bits. So 64 bits will be equal to 8 bytes. Right? That is how it is giving me the result. Now if you are interested in knowing the entire bytes, right? Entire bytes then you say n bytes will give you the total bytes occupied by the elements of the array, right, you say print a ar r r r r r r r r r r r r r r r r r r r r1.n bytes, right, and write bytes. It will be 56 bytes, right? Why? Because how many elements do we have in our arr 1 2 3 4 5 6 7, right? 7 8 are 56, right? 56 total bytes are being occupied with by ar r1. Now guys, the seventh one is as type, right? In array. This will help you change the data type of the array. Right? Change the data type of the array. Suppose I have a ar r1.d type which is uh so I say print d type, right, this is end 64, right. And now what I do is I say print uh wait let me give you a structured way print a r1, right, let me say ar r1, right, d type of array. Now a ar r2, sorry, a ar

a ar a ar a ar a ar a ar a ar a ar a ar a ar a r r r r r r r r r r r r r r r r r r r r r1 is equal to a ar r1.

As type. Right. As type. And let's say I want to convert this in np.right np.uh int32. Right. I want to create convert this in uh a ar r r int32. Right. When I do this and now when I will copy these same things you will see for yourself. Right? Now the D type was int64 and now the D type is int32. Right? If I want I can do this conversion in float also float64. Right? And then I will just copy this and I will paste it here. Right? And now you will see now the floating point has been activated. Right. It has been now activated. We can go till int we can go till int8. Right. I can go to 16. Right. And I can do this. I can go to int8 also. Yeah, like this in date also 64 32. So first one has to be Yeah. 32 whatever. Right. Like this. Okay. You can convert this also people this is later on conversion you can define the data type of the array while creation time also.

How would you do that? Suppose you are creating a ar r1 and you say np.array. Right. And suppose you take it from a list and then you just put a comma and say d type. So what will be the default data type here people? If I just do this if I just execute this what will be the default data type? And 64 is the default isn't it? But now suppose I want to change it right here. I say data type is equal to float32. Right. Sorry float64 np. sorry my bad float64. Right. And I say a ar r r11 and this will be float64 if you want float32 it will also become float32. Right. Right here while you define. Instead of using as type, you can do it right here. Right? These things will come in very handy people because you will have to save memory because when your data becomes very very big, you will be always in a crunch for memory like this. You want integers, then integers will be like this. Random.randant like this. Suppose you want to generate 0 to 6, right? And suppose you want to generate say 100 numbers like this 0 to 6 the scores 1 to 6 like this randomly. Right. Suppose you want to generate 100 scores for five different batsmen randomly it will be like this batsman number one batsman number to bat number three, four and fifth. Right. Yep. Understand the data guys. Now it's the time to understand the data. Yep.

Now guys, we have methods in numpy arrays, right? Methods in numpy arrays. So what are these methods? Now the first method we have to learn is called as reshape. Right. Reshape. Yeah. So reshape is you can use this to create you can use this to create a new shape of the array. Very very powerful guys. Very powerful. One of the most powerful methods in numpy is uh the reshape. Right. And how do we use reshape suppose we have a 1D array of 20 elements. Right. Now to reshape this reshape it we need to find the factors. Right. Factors of 20. They are what? 1 20 4 5 2 10. Right. The other factors. These are the factors. Now see what will I do. Right. Now see what will I do. Let me create a say random array. Right? Random array. I say random arr is equal to np.t random.rand. Right. And let me say I want to create it from 1 to 50. Right. And I want it to be having 20 elements. Right. So I say random arr so you will have random values inside this. Right. Random 20 values. Now seek now reshape. First I will reshape in 1+20, right? 1A 20. How will I do that? You just have to write uh print a ar r sorry sorry random.ar r random ar r.reshape.reshape and you just pass in the dimension I say 1A 20. Right. And when you will do this you will see it is coming now in 1 20 format. Right. So let me Just also write print. Right. 2D 1A 20. Right. And now what I'll do is I'll copy this and I will paste this and say 20 comma 1. Right. You will see it will be like this 20 comma 1 immediately with reshape. Right. Let me copy this let me say 2D I'm still at 2D let me say 2 10. Right. And you will see this is 2A 10. Now I can reshape it in 10 2. Right. 10 2. Right. Then I can reshape the same thing in 4A 5 and I can reshape this in 5A 4. Right. Like this guys are you able to see the power one dimension I'm able to create two dimensions and now I will take it a step further and I will write it in three dimensions. Right. How will I write it in three dimension let me say this 1 comma 2A 10. Right. This will also be three dimensions let's sorry 2A 2A 5 let me say this and now you will see I can have this in three dimensions. Right. Yeah I can also say in three dimensions like this I want to have five layers with two rows and two columns. Right. So you will have it like this also. Right. So this is How people we can reshape the array very powerfully very very powerful. Right. Very very powerful. Right. And you can take this and save print random.this and you can say print.shape. Right. So, this was the first one. Yep. Like this. Also, people if you want to visualize, we can also go this route. We can have 10 comma 2 comma 1, right? It will look like this, right? 10 layers. 10 layers you can have. Right? 10 layers you can have. Great.

Now second method which we have to learn is called as transpose. Right. Transpose method. Right? What does that do? It interchanges the dimensions like rows and columns. Right. Yes. Absolutely right. Suppose you have a matrix. Right. Which is 22 33 44 55 66 77. Right. This is a. So now when you will a transpose it. Right. The dimension. Right. Now the shape. Right. Now is 3A 2. Now this will become 2A 3. And how this will happen? Rows will now become columns and columns will now become rows. Right. So let's make column the rows. Right? Sorry columns the rows. It will be 22 44 66 33 55 77. Right? So people in transpose no information lost. It is just a change in the view. Right. Which is happening. Right? 22 44 66 33 55 77. Why do we need transposition? Suppose we have two matrix. This is 11th class mathematics. Right? One has a dimension of n cross m and the second has dimension of a cross b. If you want to multiply these two matrix say M_sub_1 and M_sub_2, right? There needs to be a satisfaction of condition. M should be equal to A. Right? M should be equal to A. Right? M should be equal to A. This should be equal to this and the resultant vector the resultant matrix which you will get will be of n crossb dimension. Right. So often times suppose this is n cross m this is n cross m. Right. m cross n and you know that m is equal to a. So what will you do? You will transpose this matrix. Right? You will transpose this matrix then it will become n cross m and then m can be equivalent to a. Right? For example, what I'm saying, we have one matrix which is 2+3 and this matrix is 2+5. Right? So, can you multiply these matrix people? Is 3 equal to 2? The answer is no. Right? The answer is no. So, what will you do? You will just transpose this and this will become 3+2 and this is 2+5. And now you can multiply this and the resultant will become 3+5 matrix. Right? So for operations like these we need transposition. Right? So how do we transpose it? How do we transpose it? So let's say again uh a r2. Right? This is a ar r2 and I want to transpose it. So I say a ar r r2 t is equal to uh np.transpose sorry a r2.transpose. Right. And now when you will see a arr you will see rows and columns have interchanged. Right. Rows and columns have interchanged with each other or let me give you one more example uh if this is not clear let me pick up This again. Right. Now I say.reshape into say 2+10. Right. So this is 2+10 and now when you want to transpose this so I'll say this E is equal to this.transpose E or A is this and you can check this now and it will be this sorry guys. So this is going to be it will be like this. Right. Transposed. And now if you want to see this. This was the original shape, right? Rows and columns have now been interchanged, transposed with each other.

Now guys, the third method, the third method which is there with us is called as flatten, right? It is called as flatten, right? What does flatten do? It reduces the dimension to one dimension. Right? Any dimension you have, it reduces it to one dimension. For example, I have this, right? And now when I say this._f, this will become this.platin. And when you will check this up, you will see that it has now become one dimension. No matter how many dimensions you have, it will become one dimension. Right? Let me take this again to show you one more example. Right? And here I say.reshape into uh uh 5+2+2. Right? I do this. My random ar r is this. Right? It has five layers, two rows and two columns. Right? So now I say this flatten is equal to this.flatten. And if you will check it now again you will see it has now flattened it out. Yep. Now why do we need this? We need this for a lot of statistical operations. We need this to feed the data into the algorithms. Right? As we will move forward you will understand the use of flattening.

Right now guys, moving on and uh as discussed, let me now show you the power of numpy. Right. Numpy over lists and other data types, right? I will not take a lot of examples. Just a second, guys. Yeah. Okay. Now guys, I told you that lists take up much more memory as compared to numpy arrays. Right? And I'm going to prove this to you. Right? Now let me use let me create a random sequence of random numbers using range in Python. Right. And say I create range of 10,000 numbers. Right. Range of 10,000 numbers so what will this give me this will give me numbers from 0 to 99999. Right. Continuous numbers. Right. So this is range I will use a range in numpy to create a similar series of numbers. Right. So let's say array is equal to np.arange. Right. Same thing same done by both. Right. I've shown you above also now let me import sis library. Right. Sis package and I will use something called as get size of. Right. Get size of what does this do get size of it's a method which calculates the bytes occupied by a single element in vanilla Python. What is vanilla Python? It is the traditional Python, right? vanilla Python. So let me just show that to you. I'll use this and I will say print. Now guys, if I get size of any random number from this range, right? Any random number. Say I get size of five, right? And I then multiply that bytes with the length of random, right? With the length of random this rand, right? Right? With the length of random, do you think I will get the bytes for the entire data structure? What am I saying is suppose uh I used range five. So what will this give me? 0 1 2 3 4. Right? This will be the output. So now I'll say get size of say I say two. Right? So suppose 2 is x and then I multiply this with the length of this series which is five. So do you think I will get 5x which will represent the number of bytes occupied by the entire data type. Anything randomly any random number this can be three. Right? Why not hurry? Why not? All of these are integers. So integers all of 64 bits assuming. So if you calculate the side of size of this and if you multiply with the total number of numbers you will get the total size isn't it? Huh? Index is in no no it's not about that. It's about the element. Right? Homogeneous elements inside this. I am saying when you use range five what is going to be the output? 0 1 2 3 4. Right. Now all these are elements elements of range. Right. All of them are elements of range. Right. Now I'm saying if I fetch the size of one element and multiply it with the length of the entire range will I get the bytes occupied by the entire range. For example. If I do this, right? If I do this, this is 28, right? 28 bytes people. 28 bytes bytes are occupied by one element of range, right? One element of range, right? Now if I just I'm saying I'm just saying if I multiply to find how many numbers range has how many numbers range has equal to 10,000. Right. So total memory occupied will be will be how much it will be 28*10,000 which will be equal to 28,000. Yeah. And how will you find this? You will say print this multiplied by length of RAM. Right. 28,000 bytes. Clear? Now yes. Now this is for the range. Now let me use another thing. So how will you calculate the length of this array? What what property and attribute will you use people? N bytes. Right? N bytes will give you total bytes occupied by the elements of the array. Right? We will use n bytes here. So I come back down and I say using n bytes for arrays. Right? And you will see what the result come. Print uh array.nbytes. Right. And I say bytes are you ready to see the result do you see what has happened how many bytes this was taking it was taking 28,000 bytes how many bytes this is taking this is taking 80,000 bytes this was taking 2 lakh 80,000 this is taking 80,000 2 lakh extra bytes of memory is taken by [Music] range. Ran is range. Yep. And if I just go to million numbers, right? Million numbers in both. See the difference it becomes, right? This is now 3 three 28 million bytes it is taking and it is taking 8 million bytes. 20 million extra bytes are occupied. Right. Now people do you believe me? Yeah that numpy wise are way more efficient in memory management as compared to the traditional data types of python. Yes. Okay. That's the first part. Now second is people performance. Right. Performance. So what I'm going to do is what I'm going to do is I am going to import times. Right. It's a it's a module in Python. Right. Suppose I say x is equal to range this much. Right. Okay. And then I have y is equal to range say this to this, right? Both of them will have equal amount of numbers. Same numbers both of them will have, right? This will have say uh 1 2 3 1 2 3 10 million values. 10 million values. This will also have 10 million values. Right? Both of them will have 10 million values. Now what I'm trying to do is I want to add them up. Right. By bit by bit I want to add them up. Right. I want to add first element of this to first element of this second of this to second of this third of this to third of this like this. Okay. I want to do this now what I'll do is I will run a counter. Right. I will run a counter which is the start time. Right. And this is given by time.time. Right. Which will give you This will give you the current time. Right? After this, I will run the operation. I will say C is equal to X+Y for X Y in zip X Y. Right? In zip X Y, right? Add X+Y bit by bit. Element by element for X and Y in zip. Zip is a function, right? Which allows you to do this operation sequentially, right? Sequentially, right? Add the elements of X and Y element by element, right? Element by element, right? Element by element, right? And then I'm going to print. Right. So start time will start. And now I will say time.time which is now the end time minus start time. So this will give me the time taken for execution. Isn't it guys? Will give me delta of time which is equal to time taken for operation seconds. Right. These many seconds will be taken. Right. So let me run this. And it takes around say 4.3 seconds, right? To do this, right? 4.3 seconds. Now guys, see what happens. You had to write this complex syntax in the traditional Python. Now let me show this on arrays. Right? What will happen on arrays? I will say a is equal to np.arange. Right. And inside a range I will pass the same values what I have taken above. Right. And I will say b is equal to np.arange and I will pass the same values inside. Right. Exactly the same now what I'll do is I will say Same thing. Right. Just I will change the execution of C will now simply become people A+B what is simple this or this this or this two. Right. Do you see the power if not I will show this to you again. Right. Later on and let me run this and you see the difference now let me just increase a couple of zeros, right? A couple of zeros. Two zeros I'm increasing in both the use cases. See, it is going on and on, right? Let's see how much time it will take to add say 2 million 1 billion numbers. 1 billion numbers I have asked my system to add and I want to see how much time it takes. Running running running. Yeah, kernel has died. Kernel has died people. I'll have to restart. Right. I will have to import numpy as np. So let me just remove one zero from both. It is taking 4 seconds removing one zero from here also. Right? And when I do this it takes 1 second. You see guys what is the difference in performance also. Right. For both of these. Yeah. And if you didn't understand this, let me give you an example. Range five. This is 5 to 10, right? And this is basically adding elements, right? Five 7 9 11 13. Right. So this will be what will be the output of this? This will be 0 1 2 3 4 and this will be output what 5 6 7 8 9. Right. So 0+5 5 6+1 7 7+2 9 8+3 11 9+4 13 and the same thing if I do here then what will happen I say 5 I say 5 and 10. Right. And I say C is equal to A+B and I say C. Same thing you get here. Right? We can move on. Right? The next bit guys which we have to understand the next bit which we have to understand is called as the indexing in numpy arrays. Right. Indexing in numpy arrays. Right? How do we index the elements? Right? How do we index the elements? Indexing in num py arrays. Right. Indexing in num py arrays so again you know indexing from basic python so let's start with 1D for 1 arrays. Right. I will use ar r r1 yeah this is ar r r1 now. Right. This is a ar r1 Okay. Now people what I want to do is what I want to do is I want to fetch. Right. I want to fetch. Right. You can slice and dice let's say dice 33. Right. So again as for a normal indexing of list what is 33 what is the index of 33 people so you will say the same thing print. Right. A R R1 squared bracket 2 and you will get 33 for yourself. Right? If you wish to slice same things, right? 33 to say 66. What is the index? 33 is 2. 2. Which one? 3 4 5 and six. Right? We will write 3 to 6. Not five. Hurry. Five is not included. Remember? So print a ar r r1 2 is to 6 and you will get 33 44 55 66. Right. Simple indexing. Please try it out. Please try it out and if you want you can have this code also. You can write this code. You will always have clarity that why do we use it. Moving on people. Moving on. Let's see indexing in a 2D array. Right? And before I explain this

To you, let me take you here. Right? So a 2D array will be what? Right? This is a 2D array. It has three rows and three columns. Right? Rows. Columns. So now for rows, indexing will start from zero. So if you have to fetch this particular row, right, so what will be the index? It will be row 0. If you have to fetch this particular row, the index will be one. And if you have to fetch this particular row, the index will be two. Similarly, for columns, if you have to fetch this particular column, right, you will have column is equal to zero. This particular column, column equal to 1. This particular column, column equal to 2. Right? Indexing will start from minus 1 again, right? 0 1 2. So let me just show that to you, right? Let's say I call arr2; this is my arr2. Now, indexing first row, right? How will I do that? I will say print arr2, and I will write how will I write this? I will write row 0, right? Row 0. So what will this give me? This will give me this. Right? The syntax is row space column. Right? So now if you just pass one, it will give you only rows, right? Row 1, row 2, row 3, right? Similarly, if you want to create it for columns, right? What will you say? Sorry. Uh, what [Music] columns? Uh, it was zero. Column 1, column 2, column 3. Right. Right. So what is my column 1? 76 89 98 76 89 98. Right. And if you want column 2, uh, sorry, if you want column 2, this is the column 2, and this is column 3. Right, guys? 90 999 999 independently.

Now, if I ask you to fetch me a particular element, element, right? Say I want you to fetch me 78, right? How will you fetch 78? You will say print arr2. What is the row for 78? Which row does it belong to? 0 1 2. To which row does 78 belong? One row. Which column it belongs to? 0 1 2. 1. Hurry, check 1 0 whether it belongs to zero column, first column, or second column, right? And you will get uh, sorry, 0 1. So this is two, right? 78, right? Second row, first column, right? Second row, first column. How to define it as an? Right? This is your matrix. Right? Now, what are the index positions for this? This row is zero. Row, first row, second row. This is your zeroth column, first column, second column. Yeah. Yes. No. Maybe. Are we understanding this? This much is clear, the indexing of rows and columns. Now, now if you have two, so the syntax is, syntax is, say this is matrix A. So you will say A, in this A matrix, you will write row, comma, column, right? Suppose I want to access only zeroth row. So there will be no column. So what will you get? Row number zero, right? Row number zero. And you will just put it like this. Or you can put a comma and put colon. Colon means what? Take everything. So I want all three columns together. Zero row and all three columns. So this will be your show me. Let me show that to you. Right? See this zero row and all the columns. So what will be your answer? 76 88 90. Right? Then if you want to access the second row, 89 90 99 89 90 99, one and like this, clear.

Now, similarly, for columns, what will happen? You will take all the rows, comma, which column do you want? If you say two, what will be the result for this? What numbers will you get for this? Colon, comma, 2, you will get 33 66 999. Now, coming to the element, right? Suppose now you want to fetch 55, right? So what will you write? You will write which row does it belong? Follow this syntax. It belongs to the first row. Which column does it belong to? First column. So what will you get? 55. Come to, come to this example. Now you want 78 in this particular array or matrix, right? Where is 78? Which row does it belong to? Is it first row? Check carefully. Zero row, first row, second row. Yeah. So I put two here. Now which column does it belong to? First column, second column. Sorry, zero column, first column, second column. It belongs to the first column. So I put one here. So when you put this syntax, you will get 78. Clear?

Now, the third thing is slicing through the array. Right? Now, for this, I want, I want 90 99 78 99. That means what? I want this, this, and this. Let me take you to the PPT first. Now, what I'm asking you to fetch me? I'm asking you to fetch me these four numbers. Right? These four numbers. So here, what will you write? Which rows are included in this, people? Which rows are included in this? Row one to all. Right? One to all. Yeah. Not two. One to all. Right. And you leave everything like this. If you have the last row, you leave it empty after the colon. Comma. Which columns do you want for this? One and two. Right? And when these will intersect, when these will intersect, you will get this area. You will get this shaded area, green shaded area. So I will say I need from column one to all the columns. Right? So what will this fetch you? What will this fetch you? This will fetch you 44, 55, 66, and 77, 88, 99. Right? What will this fetch you? This will fetch you 22, 55, 88, 33, 66, 99. What are the commonalities between both of these access, both of these slices? What are the commonalities? It is only this much, right? 55, 66, 88, 99, right? So, do you think you will get your result? Yeah. Let's check it here. Right? I say print arr2. Right? And this I say I need row zero, sorry, row one to empty, and then column 1 empty, and you will get 90 99 8 indexing. Yeah.

Now, guys, let's check it for three dimensions. Right. 3D. I say arr3, right? This is my arr3, right? So let me just put an example here. Suppose my arrays are 1 1 22 33 4 55 66 7 88 91. Right? This is my first. Then, behind this, I have another matrix which is uh, 111 222 333 444 555 666 777 888 9999. Right. And in my third one, I have 1 1 222 33 33 3 4 44 44 44 44 44 44 44 44 44 44 44 44 44 44 44 44 44 44 44 44 44 4 55 555 66 66 666 7 88 8 9. Is that, is it qualifying for a 3D array? Can you access anything on this particular array if given a chance separately? Like this one. Now, people, if you want to move between layers, right? If you want to move between layers, do you think I have told you one particular access which is what? Which is the layers. So what number will be given to this layer? Layer number zero, layer number one, and layer number two. So now if you have to access this 555, which layer you will go to first? You will go to first layer, then which row will you go to? You will go to first row and first column. So if you pass this syntax, what will you get? You will get 555, right? Problem solved. The only bottleneck was the layers part, and you have an additional parameter or argument for this layer. So if I come back to my example, and if you have to access this 555, right? How will you do that? I say print arr3, right? And this I say, and in this I say what? I have to access this 555. Which, which layer number is this? This is layer number zero. Right? This is layer number zero. And this is layer number one. So I say layer number one. Which row is this in this particular layer? It is layer number. Sorry, it is row number one and column number one. And when you do this, you will get 555. Right? If you want 999, what you will do? You will change this to 2, 2, right? If you want 777, or let's say if you want 98, what will you do for 98? Layer number zero, row number two, column number zero, then you will get this 98. Clear, guys? Yeah. Layer, row, column. In the same way, you will slice it. Right. You will slice it. Yep. Right. I'll write the syntax so that you don't get confused. It is array square bracket layer row column. Right? Layer row column. Right. This is the syntax. Perfect. Right. Great.

We can also perform some operations, right? Plus, minus, multiplication, division between two arrays very, very easily, right? It should not pose any problem to us, right? We can do all the operations which we want to, right? Between two arrays, right? For example, right? You had, you have, right? Operations on parents. So let's say uh, let's see as a list, right? List. So we have list is equal to 1 2 3 4 5. Right? Now, suppose you want to square each element of the list. Right? What will you do? You will say sql is equal to x to the power of 2 for x in l, right? And then you will say print sql, and you will get the squared of list. Right? Original list, squared list. Now, in array, what will happen? Suppose I say a1 is equal to np.array and l. I create the same array out of this list. I say print original array, right? I say a1, right? This is my original array, same as this. Now I want to square it, square the elements of arrays. Very, very simple, nothing you require. You just say, you just say sqa1 is equal to, sqa1 is equal to a1 to the power of two, right? A1 to the power of two, right? And then you print the, right? You get the squared R. Suppose you want to find the mean of [Music] the mean of numbers using list. Right? So what will you do? You will say mean is equal to sum of l, right? List divided by length of list, right? And you will get the means, right? Which is three for this one, right? This original list. Now let me show you in arrays, right? How will you do this? You will say mean is equal to np.mean, and you will just pass a1, right? And when you will check mean, you will get 3.0. Right? Nothing like this, direct, right? Direct like this, right? You can, you have, I've already showed you add, I've already showed you uh, square, and then I believe you can understand that what all operations are possible using the arrays, right? Leveraging the power of arrays, right? Also, guys, in the arrays, right? In the arrays, what you can do is you can perform, you can perform some string operations, right? Very powerful string operations, right? So, for example, let's say I have a, I have an array, right? I have an array of names of people, right? So I say names is equal to np.array, right? And I say Radha, uh, then I say D, and then I say Maduk, right? These three names I had, right? So you can check the names; they will be like in the array, right? And the data type will be U6, which is a representation of strings, right? Now, guys, suppose you want to capitalize, you wish to capitalize the names, right? Of people. What will you do? You will say print me np.char, right? np.char.capitalize, right? Capitalize, and inside this you will pass names, and you will see all the names have been capitalized, right? You see this R has been capitalized, D has been capitalized, M has been capitalized, right? You can convert them into upper if you want. Print np.char.upper names, and you will have all of them in caps lock. You can say print np.char.lower, you will have them in lower, right? You can put the title, right? Suppose I say uh, title is equal to np.array, and I say Raghav Goyal, right? Then I say Dhapati, right? And I say Madras, right? I say these three things. Now if I say print np.char.title, right? And I say titles, you will see that all the words will be in capitalized mode. Raghav Goyal, D and S of Dhapati, M and V of MaduMa are now capitalized. I can also replace something if I wish to. Suppose I want to replace Madu with say Suri, right? I will say print, right? np.char.replace, right? And you will say where you want to replace. I say title, and in this I want to replace Madu with Suri, right? And you will say it will be Suri, right? Madu has been replaced with Suri. Right. Right. If you want to calculate the characters of strings, right? You can do that. Print np.char.str_length of titles, and you will get 11 characters are there in here, 15 are there here, and 11 are here in this particular thing. Right? You can do much more powerful things also. Let me show you one complex function. Right? Suppose I have fname is equal to np.array. Right? And we have Radh Madu, right? And we have lname ari warm, right? We have these two things. Now I can create a new array fullname by simply, right? By simply saying np.char.add, right? And I say uh, fname, right? Comma lname, right? Plus. Yep. Fullname. Yeah. Ra. I just was trying to add a space in between. Anyway, right. We can do that, right? You can also, people, search in arrays, right? Very powerful. Again, search in arrays using where, right? Right. You can say suppose a2 is equal to np.array, right? And I'll say 1 1 22 33 3 4 4 5 66, right? And now you have to say a is equal to np.where, right? And in this you say a2 greater than 20, right? And when you will check a, uh, it's giving me the index. Why is it giving me the index, or is it giving me the index? Does it always return index? One more thing is you can find a, can find an a letter, right? Through a letter. You can find an element through a letter. Guys, these are all some tricks which you should know because you will be dealing with data and you need to pull data, right? You need to pull data a lot, right? Based on conditions and based on things. Suppose you want to find out the names which have G in them, right? Or RA in them. So how will you do this? I will say print, right? And I will say uh, np.char.find, right? And I will say find this inside fullnames, right? And find me ra a, right? Ra a two p it uh, right? It returns true because it has found it here, right? So I don't want to tell you indexing through this, but anyway, you should know this, just, just assume this that I'm telling you to write this, okay? Because this is much easier when we will go to pandas, right? Just uh, write it as a syntax, okay? Greater than equal to zero. I hope this is clear.

Hello everyone. I am and welcome to today's video where we will be talking about LLM benchmarks, tools used to test and measure how well large language models like GPT and Google Gemini perform. If you have ever wondered how AI models are evaluated, this video will explain it in simple terms. LLM benchmarks are used to check how good these models are at tasks like coding, answering questions, and translating languages or summarizing text. These tests use sample data and a specific measurement to see how well the model perform. For example, the model might be tested with a few examples like few-shot learning or none at all like zero-shot learning to see how it handles new tasks. So now the question arises, why are these benchmarks important? They help developers understand where a model is strong and where it needs improvement. They also make it easier to compare different models, helping people choose the best one for their needs. However, LLM benchmarks do have some limits. They don't always predict how well a model will work in real-world situations, and sometimes models can overfit, meaning they perform well on test data but struggle in practical use. We will also cover how LLM leaderboards rank different models based on their benchmark scores, giving us a clear picture of which models are performing the best. So stay tuned as we dive into how LLM benchmarks work and why they are so important for advancing AI. So without any further ado, let's get started. So what are LLM benchmarks? LLM benchmarks are standardized tools used to evaluate the performance of large language models. They provide a structured way to test LLMs on a specific task or question using sample data and predefined metrics to measure their capabilities. These benchmarks assess various skills such as coding, common sense reasoning, and NLP tasks like machine translation, question answering, and text summarization. The importance of LLM benchmarks lies in their role in advancing model development. They track the progress of an LLM, offering quantitative insights into where the model performs well and where improvement is needed. This feedback is crucial for guiding the fine-tuning process, allowing researchers and developers to enhance model performance. Additionally, benchmarks offer an objective comparison between different LLMs, helping developers and organizations choose the best model for their needs. So, how LLM benchmarks work? LLM benchmarks follow a clear and systematic process. They present a task for the LLM to complete, evaluate its performance using a specific metric, and assign a score based on how well the model performs. So here is a breakdown of how this process works. The first one is setup. LLM benchmarks come with pre-prepared sample data, including coding challenges, long documents, math problems, and real-world conversations. The tasks span various areas like common sense reasoning, problem solving, question answering, summary generation, and translation, all presented to the model at the start of testing. The second step is testing. The model is tested in one of three ways: few-shot. The LLM is provided with a few examples before being prompted to complete a task, demonstrating its ability to learn from limited data. The second one is zero-shot. The model is asked to perform a task without any prior examples, testing its ability to understand new concepts and adapt to unfamiliar scenarios. The third one is fine-tuned. The model is trained on a data set similar to the one used in the benchmark, aiming to enhance its performance on the specific task involved. The third step is scoring. So after completing the task, the benchmark compares the model's output with the expected answer and generates a score, typically ranging from zero to 100, reflecting how accurately the LLM performs. So now let's move forwards. Let's see key metrics for benchmarking LLMs. So, LLM benchmarks use various metrics to assess the performance of large language models. So here are some commonly used matrices. The first one is accuracy or precision. It measures the percentage of correct predictions made by the model. The second one is recall, also known as sensitivity. It measures the number of true positives, reflecting the correct predictions made by the model. The third one is F1 score. It combines both accuracy and recall into a single metric, weighing them equally to address any false positives or negatives. F1 score ranges from zero to one, where one indicates perfect precision and recall. The fourth one is exact match. It tracks the percentage of predictions that exactly match the correct answer, which is especially used for tasks like translation and question answering. The fifth one is perplexity. Here it will tell you how well a model predicts the next word or token. A lower perplexity score indicates better task comprehension by the model. The sixth one is BLEU, bilingual evaluation understudy, is used for evaluating machine translation by comparing n-grams, sequences of adjacent text elements, between the model's output and the human-produced translation. So these quantitative metrics are often combined for more thorough evaluation. So in addition, human evaluation introduces qualitative factors like coherence, relevance, and semantic meaning, providing a nuanced assessment. However, human evaluation can be time-consuming and subjective, making a balance between quantitative and qualitative measures important for comprehensive evaluation. So now let's move forward and see some limitations of LLM benchmarking. While LLM benchmarks are valuable for assessing model performance, they have several limitations that prevent them from fully predicting real-world effectiveness. So here are a few. The first one is bounded scoring. Once a model achieves the highest possible scores on the benchmark, that benchmark loses its utility and must be updated with more challenging tasks to remain a meaningful assessment tool. The second one is broad data sets. LLM benchmarks often rely on sample data from diverse subjects and tasks. So this wide scope may not effectively evaluate a model's performance in edge cases, specialized fields, or specific use cases where more tailored data would be needed. The third one is finite assessment. Benchmarks only test a model's current skills, and as LLMs evolve and new capabilities emerge, new benchmarks must be created to measure these advancements. The fourth one is overfitting. So if an LLM is trained on the same data used for benchmarking, it can lead to overfitting, where the model performs well on the test data but struggles with real tasks. So this results in...

Scores that don't truly represent the model's broader capabilities. So now, what are LLM leaderboards?

So LLM leaderboards publish a ranking of LLMs based on a variety of benchmarks. Leaderboards provide a way to keep track of many LLMs and compare their performance. LLM leaderboards are especially beneficial in making decisions on which models to use. So here are some.

So in this, you can see here OpenAI is leading, and GPT-4 is second, and the Llama is third with 405 parameter B, and 3.5 is there. So this is best in multitask reasoning. What about the best in coding?

So here OpenAI 01 is leading. I guess this is the Orca one, and the second one is 3.5 Sonnet, and after that, in the third position, there is GPT-4. So this is in best encoding.

So next comes fastest and most affordable models. So fastest models are Llama 8B parameter, 8B parameter, and the second one is LLaMA 70B, and the third one is 1.5 flash. This is Gemini 1 and lowest latency, and here it is leading Llama again in cheapest models. Again, Llama 8B is leading, and in the second number, we have Gemini flash 1.5, and in third, we have GPT-4 mini.

Moving forward, let's see standard benchmarks between Claude 3, Opus, and GPT-4. So in general, they are equal in reasoning. Claude 3, Opus is leading, and in coding, GPT-4 is leading. In math, again, GPT-4 is leading. In tool use, Claude 3, Opus is leading, and in multilingual, Claude 3, Opus is leading.

So let's start with the data science interview questions and answers, and the number one problem we will be facing is real-world problem-solving, and the question one is handling missing data in predictive modeling.

So imagine you have given a data set where 30% of the data for a key predictive variable is missing. This variable is crucial for a predictive model. How would you handle this situation to ensure the integrity and performance of your model? And please describe your approach step by step.

So starting with the answer, you can start with handling missing data sets. This is a common challenge in data science, and it's important to address it carefully to maintain the accuracy of your model. And here's how you could approach this situation. The number one point could be identifying the missing data.

So first, you need to understand where the missing values are in your data set. You can do this by using a simple code in Python with libraries like pandas. For example, you can use the data.isnull().sum() function, that will show you the count of missing values in each column. Then you can analyze the pattern. Determine if there's a pattern to the missing data. Is it random, or is it missing for a reason? This can affect your approach. If the data is missing at random, the methods you use might be different than if the data is missing systematically.

So, choosing a method for imputation. Let's see the next method that is choosing a method for imputation. So, if the missing data is numeric, you might replace missing values with the mean or median of that column. This is simple and effective but can be used primarily when the data is missing completely at random. Then comes model-based imputation. Sometimes you can use other variables in the data to predict missing values using a regression model. This can be more accurate but is also more complex. Then we'll use the k-nearest neighbors algorithm. But before that, we have a code snippet here that could be used for the implementation of imputation. You could use Python or R.

And now moving on, we'll see the k-nearest neighbors algorithm. So this method predicts the missing values based on how closely related the data points are to each other. So after imputation, it's crucial to check how your changes have affected the overall data set and model performance. Sometimes filling in too many missing values can introduce bias. And then we have visualization. To help understand before and after the imputation, you could visualize the distribution of the variable using histograms or box plots. This helps in seeing how the imputation has changed the statistical properties of the data, and by following these steps, you can handle missing data thoughtfully and maintain the integrity of your predictive model.

Now moving to question number two, that is based on evaluating model overfitting. So the question is: you have developed a predictive model, but you suspect it might be overfitting the training data. How would you test and address the issue? Please explain your steps and the techniques you would use.

So you could start the answer by explaining what is overfitting. So overfitting is a common problem where a model performs well on training data but poorly on unseen data, indicating it's too closely fitted to the training data's specific details and noise. So now we'll see a step-by-step guide on how to address this. The number one step is cross-validation.

So one effective way to test for overfitting is by using the cross-validation technique. Cross-validation involves splitting your training data into multiple smaller sets, that is folds, and then training a model on some of these sets and validating it on the others. So this helps you understand if the model's good performance is consistent across different subsets of data. For example, in Python, you can use the cross_val_score function from sklearn.model_selection. So this is the code, and this is the code snippet of Python that you can use for the cross-validation, and here we are importing from sklearn, that is the module, and we're importing cross_val_score, and here we have used the cross_val_score function, and then we have printed the average cross-validation score.

And the next step we will do is running the cross-validation model. So this is your predictive model that you have already built using scikit-learn. And here's the x_train. These are the x input features of your training data. And y_train, these are the output labels of training data. So we are running the cross-validation model here. This is your predictive model that you have already built using scikit-learn. So x_train here, that means these are the input features of your training data, and y_train here means these are the output labels of training data, and cv=5. This parameter tells the function to split the data into five parts, that is folds. And the model is trained on four of these parts, and the remaining part is used for testing. So this process rotates until each part has been used for testing once, and the printing results, that is score.mean(). So this calculates the average of the scores obtained from each cross-validation. This average score gives you an idea of how well your model is likely to perform on unseen data. A consistent score across different folds suggests your model is generalizing well rather than overfitting to the training data.

So now moving to the next point, that is training versus validation error. So plot the training and validation errors as a function of training epochs or complexity of the model. A model that overfits will show a low error on training data and a high error on validation data as it trains further. Then we have pruning the model. If you confirm that the model is overfitting, consider simplifying it. This might mean reducing the number of parameters by selecting fewer features, using regularization techniques like lasso or ridge, or choosing a less complex model. After this step, we will move to the regularization technique step.

So these techniques add a penalty to the loss function used to train the model, which can discourage complex models that overfit. Then we have common methods that include L1, that is lasso, and L2 ridge regularization. And here's how you can add L2 regularization in Python. So this is the code snippet here. And what we have done here is we are creating the ridge model, and we have applied alpha=1.0. So this parameter controls the strength of the regularization. A higher alpha value increases the regularization effect, which helps reduce model complexity and combat overfitting. The alpha value can be tuned to find the optimal balance between bias and variance. And now coming for the fitting the model. So model.fit, and in that we have X_train and Y_train, that trains the ridge model on the training data. It adjusts the weight of the features in X_train to predict the Y_train while also considering the regularization term. This helps prevent the model from fitting too closely to the noisy aspects of the training data. And then we are re-evaluating the model. After making adjustments, it's important to re-evaluate the model again using the same cross-validation technique to see if the issue of overfitting has improved. And then we have visualization. To help illustrate overfitting, you could create a plot showing the training and validation errors over the number of epochs or model complexity. So by using these techniques, you can identify if your model is overfitting and take steps to correct it, ensuring it performs well not only on the training data but also on new unseen data.

So now move to the next question, that is question number three, and it is based on real-time data stream processing. And the question is: you are tasked with building a model to predict stock prices in real time. The data comes in every second, and you need to update your predictions accordingly. Describe how you would set up your system to handle this type of data effectively, and what tools and techniques would you use and why.

So you could start answering this question with handling real-time data. So handling real-time data, especially for something as volatile and fast-paced as stock prices, requires a robust system that can process and analyze data quickly and accurately. So here's how you could approach this. We will set up such a system, and we'll have some steps. So starting with the steps. So the first step is choosing the right tools. The right tool would be Apache Kafka. So this is a popular tool for handling real-time data streams because it allows you to publish and subscribe to streams of records, that is data, and it can handle high throughput with low latency. Kafka acts as a buffer and manages the flow of data, ensuring that your system doesn't get overwhelmed, and you can also use Apache Spark, especially Spark Streaming. This is excellent for processing the data. It can process data in real time and perform complex operations like windowing, grouping data into chunks of a specified time period, and aggregating, that is summarizing data. So you can modify it and perform the predicting of stock prices, and then the step is the data processing pipeline, and the first step comes here is injection. Data first enters the system, typically through Kafka, which collects data sent from the stock market, and then we do the processing. So Spark Streaming takes over here. Here you can apply transformations and run your predictive models on the data. For example, you might calculate moving averages or other indicators that feed into your stock price prediction model. And then comes the output. Finally, the predictions are outputted. This could be to a dashboard for traders, an automated trading system, or even stored for further analysis. And then we develop the model. Now comes the model development. You would likely use a machine learning model that can update quickly and incorporate new data as it arrives. Models such as ARIMA for time series forecasting or more complex machine learning models like recurrent neural networks (RNNs) can be suitable. The model should be retrained or fine-tuned periodically with new data to ensure it stays accurate. Now we'll come to scalability and reliability. So ensure your system can scale as data volume increases. This might mean adding more servers or optimizing your data processing code. Implement monitoring to catch any issues early, like delays in data processing or model performance drops. And now we'll see the step that is visualization and monitoring. Consider setting up a real-time dashboard that shows key metrics like prediction accuracy and processing time. This helps in quickly spotting when something goes wrong. By setting up your system with these tools and strategies, you can effectively handle the challenge of predicting stock prices in real time.

So now we move to the next question, that is question number four, and this will be based on scalable data analytics. So we have covered two questions that were a bit code-based questions, and now we'll see other questions that would be based on scalable data analytics, or they might be on different areas, and with the 13th question, we'll start again with the coding ones.

So moving with question four, that is based on scalable data analytics, and the question is: given a scenario where your organization suddenly needs to scale its data analysis capabilities due to an influx of data that would be 10 times the normal volume, how would you handle this situation to ensure your data analytics processes remain efficient and accurate? What technologies would you consider, and what steps would you take?

So you can start answering this question with handling a sudden increase in data volume. This requires a strategic approach to scaling your analytics infrastructure without compromising on efficiency or accuracy. So we'll see some steps that you could effectively manage this scenario that you would start answering the interviewer that we can start by evaluating the current infrastructure's ability to handle increased loads. This includes assessing your databases, servers, and analytical tools to identify potential bottlenecks or limitations. Then you could move to the next step that would be choosing scalable technologies to manage the increased data volume. Consider leveraging cloud-based solutions such as Amazon Web Services, Google Cloud Platform, or Microsoft Azure. These platforms offer scalable resources, which can be adjusted accordingly to the data load, ensuring you only pay for what you use. Integrate big data technologies like Apache Hadoop for distributed storage and Apache Spark for fast data processing. These tools are designed to handle massive volumes of data efficiently and can scale up to meet significantly increased demands. Now we move to the next step that would be optimizing data processing. So implement data partitioning and indexing strategies to improve the efficiency of data queries. This will help in managing large data sets by breaking them into smaller, manageable chunks and speeding up search operations, and use real-time data processing frameworks like Apache Kafka or Apache Flink, which can handle high throughput and provide timely insights from large data streams. And the next step would be automation and monitoring. Automate routine data processing tasks to reduce the manual effort and speed up the analysis. This can be done through scripting or using workflow automation tools. Set up comprehensive monitoring systems to track the performance of your data processes. Tools like Prometheus for system monitoring and Grafana for analytics and monitoring dashboards are useful here. They help ensure that the system is running smoothly and alert you to potential issues before they become critical. And the next step will be regular evaluation and scaling. Continuously evaluate the performance of your analytics infrastructure. As your data grows, keep adjusting and scaling your resources to maintain optimal performance. Plan for periodic reviews of your technology stack and infrastructure to ensure they remain aligned with your data needs and organizational goals. By following these steps, you can ensure that your data analytics processes are prepared to handle sudden surges in data volume effectively, maintaining the integrity and speed of insights. So this was all for question four.

Now moving to question five, and this is based on integrating machine learning models into production. And the question is: you have developed a machine learning model that performs well in a testing environment. Now you need to integrate it into your production environment where it will be used in real-time applications. What steps would you take to ensure the successful deployment and operations of the model in production?

So we'll start answering this by successfully deploying a machine learning model into production. This involves several critical steps to ensure it performs as well in real-time operations as it does in testing. So you would have a clear pathway to make the interviewer understand. We will start with the pathway with the first step that would be model validation. So before moving anything into production, re-validate your model's performance using a separate validation data set. This helps confirm that the model generalizes well to new, unseen data. The next step will be preparing the production environment. Ensure that the production environment is ready to handle the model. This includes setting up the necessary hardware and software, ensuring that it can handle the expected load and that all dependencies are correctly installed and configured. Then the next step comes that is model wrapping. Wrap your model in an API, that is application programming interface, making it accessible to other parts of your software infrastructure. Frameworks like Flask for Python can be used to create a simple web server that listens for data inputs and provides model outputs. Then comes the next step that is deployment strategies. Consider using containerization tools like Docker, which can help encapsulate your model and its environment, ensuring that it works uniformly across different development and production settings. And then we'll use deployment strategies like blue-green deployment or canary releases to minimize downtime and reduce the risk of introducing a faulty model into production. And then comes the next step that is monitoring and logging. Implement logging and monitoring to track the model's performance and health in real time. Tools like Prometheus for monitoring and ELK (Elasticsearch, Logstash, Kibana) for logging help in quickly identifying and diagnosing issues in production. And then comes the next step that is performance tuning. Monitor the model's performance over time. If the model's performance degrades, or if new data shows different patterns, you may need to retrain or fine-tune the model to maintain accuracy. And after this step, there's a step for the feedback loop. Set up a feedback loop where predictions and outcomes can be compared. This feedback is crucial for continuously improving the model and catching any drift in data or changes in external conditions that affect the model. And after this comes a last step that is legal and compliance checks. Ensure all the data used by the model in production complies with privacy laws and regulations. This is crucial for maintaining trust and legality, especially when handling sensitive information. So by carefully planning and executing these steps, you can smoothly transition your machine learning model from a testing environment to a fully functional component of a production system. So this was all about question number five.

Now moving to question number six, that would be based on data-driven decision-making. And the question is: your company wants to shift towards more data-driven decision-making. You have been tasked with developing a strategy to implement this. What steps would you take to ensure that the data at all levels of the organization is utilized effectively to make informed decisions, and what challenges might you face and how would you address them?

So you can start answering this by implementing a data-driven decision-making strategy. This will require a comprehensive approach to ensure that reliable data is accessible and effectively used across all levels of the organization. And now we can develop and deploy this strategy. And similarly, you could tell this strategy to the interviewer. So the number one step will be assessing the current data infrastructure. Start by evaluating the existing data infrastructure to understand what data is available, how it is stored, and how it is currently used. This assessment will help identify gaps in data collection, storage, and access that need to be addressed. Now we move to the next step that is developing a data governance framework. Implement a data governance framework that defines who can access data, how it can be used, and who is responsible for maintaining its quality. This framework ensures data integrity and security, which are critical for making reliable decisions. Now we move to the next step that is training and empowerment. So train employees at all levels on the importance of data-driven decision-making and provide them with the tools and knowledge necessary to analyze and interpret data. This might include training sessions, workshops, and ongoing support to ensure everyone can use data effectively. Now moving to the next step that is implementing analytical tools. So deploy user-friendly analytical tools that can integrate seamlessly into the daily workflows of employees. Tools like Tableau, Microsoft Power BI, or even advanced Excel techniques can provide powerful data analysis capabilities without requiring extensive technical knowledge. After this, we'll move to the step that would be creating a centralized data platform. Develop a centralized data platform where all organizational data can be accessed and analyzed. This platform should be scalable and secure, providing a single source of truth for the organization. And then we have promoting a data-driven culture. So foster a culture that values data-driven decision-making; encourage experimentation and learning from data-driven initiatives. Celebrate successes and learn from failures to continually improve the use of data-driven decision-making, and there would be some challenges and solutions for that. So one major challenge we know here is resistance to change, as some employees may prefer traditional decision-making methods. So address this by demonstrating the tangible benefits of data-driven decisions through pilot projects and success stories. So data silos can also hinder effective data use; promote cross-department collaboration and integrate disparate data sources to overcome this challenge. After that, you can monitor and do continuous improvement. So by systematically implementing these steps, you can transform your organization into one that leverages data at all levels to make informed and effective decisions. And after answering in these steps, you could make the interviewer have a truth and a faith in you that you could make these models.

Now move to the next question, that is question number seven, and that is based on handling large data sets. And the question is: your project involves analyzing extremely large data sets, potentially exceeding terabytes in

size. What strategies would you use to manage and analyze such large data sets effectively? Describe the tools and techniques you might employ. You could start this with answering that working with large data sets, especially those in the terabyte range, presents unique challenges in terms of storage, processing, and analysis. So we'll have a structured approach to handle these challenges effectively.

We'll start with the data storage; that would be using distributed file systems. Consider using a distributed file system like Hadoop Distributed File System (HDFS) or Amazon S3. These systems are designed to store vast amounts of data across many servers, offering high availability and fault tolerance.

And then comes the next step, that is data processing. Leverage big data processing frameworks. Tools like Apache Spark are ideal for processing large data sets because they handle distributed computing effectively. Spark can perform data processing tasks much faster than traditional disk-based processing due to its in-memory computing capabilities.

And next, we could start with efficient data sampling. There are many sampling techniques that we can use. When the data set is too large to handle even with powerful tools, consider using data sampling techniques to reduce the size to a manageable level without losing significant insights. Ensure that the sample represents the whole data set accurately.

And then comes optimization of data queries: indexing and partitioning. Optimize your data queries by implementing indexing and partitioning. This can drastically reduce the time it takes to perform queries by limiting the amount of data scanned.

And then we can do scalable analytics. We'll move to the next step, that is scalable analytics. In that, we could start with parallel computing. Use parallel computing capabilities of frameworks like Spark or Dask to analyze data across multiple nodes. This helps in scaling up your analytics operations to handle large data sets effectively.

Now we'll move to cloud-based analytical tools. Consider using cloud services like Google BigQuery or AWS Redshift, which are designed to handle massive data sets and complex analytics with ease.

After this step, we'll move to data cleaning and pre-processing. Here we will automate pre-processing tasks. We'll use automated tools to clean and pre-process data. This includes handling missing values, normalizing data, and removing duplicates, which can be particularly challenging with large data sets.

After this step, we'll move to the step that will visualize large data sets. So we'll use specialized tools. Those tools could be Tableau or Power BI that can handle large data sets by aggregating data and using efficient backend technologies. For more detailed exploration, tools like Plotly or Bokeh can be used as they offer capabilities to interactively visualize large volumes of data.

And after that, there would be a step for regular maintenance and updates. That could be continuously monitoring the data quality. As new data comes in, you can continuously monitor its quality.

After this step, you could integrate all these strategies and tools into your workflow. You can effectively manage and extract valuable insights from extremely large data sets, thereby supporting robust data-driven decision-making. You could answer the whole strategy to the interviewer.

Now moving to question number eight, that is based on optimizing machine learning models. The question is: during model development, you have noticed that your machine learning model is underperforming. What steps would you take to diagnose the problem and optimize the model's performance? What techniques and tools would you use?

You can answer this by starting with optimizing a machine learning model that is underperforming. Optimizing a machine learning model that is underperforming involves several steps to diagnose and improve its accuracy and efficiency. Here we will have a structured approach to tackle this issue. You could start this with the number one step, that is diagnosing the problem.

Evaluate model metrics. Start by thoroughly evaluating the performance metrics of your model. For classification tasks, look at accuracy, precision, recall, and the F1 score. For regression tasks, consider R-squared, mean squared error (MSE), and mean absolute error (MAE).

Then you can move to the next step, that is use plots like ROC curves for classification models and residual plots for regression to visually assess where the model is going wrong.

After that, we'll move to the next step, that is data quality and quantity check. Inspect the data. Sometimes the quality and quantity of data can be the root cause of poor model performance. Ensure the data is clean, well pre-processed, and sufficient. Look for issues like missing values, outliers, or imbalanced classes.

After this, we'll move to the feature engineering step. Experiment with creating new features or transforming existing ones to provide better predictive power.

Then we have the next step, that is model tuning and configuration. After feature engineering, we'll move to the next step, that is model tuning and configuration. So, hyperparameter tuning. Use techniques like grid search or random search to find the optimal settings for your model's parameters. Tools like scikit-learn's GridSearchCV or RandomizedSearchCV can automate this process.

Implement cross-validation to ensure that the model's performance is consistent across different subsets of the data set.

Then we have the next step, that is trying different models. Experiment with algorithms. If initial models are underperforming, try different algorithms that might be better suited for the problem. For instance, if you started with linear regression and it's not performing well, consider more complex models like random forest or gradient boosting machines.

After this, we have ensemble methods. Use techniques like bagging, boosting, or stacking to combine the predictions of multiple models to improve overall performance.

After this step, we have feature selection, which includes reduce dimensionality. Use techniques like principal component analysis (PCA) to reduce the number of features, which might help in improving model performance by removing noise and redundancy.

Select important features. Use model-based techniques to identify and keep only the most important features that impact the outcome.

Then comes the last step, that is regular updates and retraining. Continuously monitor the model's performance over time. As new data becomes available, update and retrain the model to adapt to any changes in underlying patterns.

After that, you could have a consultation and collaboration. Work with other teams. By methodically addressing each of these areas, you can diagnose why your machine learning model is underperforming and take steps to optimize its accuracy and efficiency.

So this was all about question number eight. So let's start with question number nine, and this is based on handling unstructured data. The question is: you are given a large amount of unstructured data, including text, images, and videos. What strategies would you use to manage and analyze this type of data effectively? Describe the tools and techniques you might employ.

You can start answering this question by describing that dealing with unstructured data can be challenging due to its lack of predefined format or structure. However, with the right strategies and tools, you can effectively manage and analyze it to extract valuable insights. There will be an approach on how you can do that. So we will discuss the approach here, starting with the steps.

The number one step will be data categorization and organization. In this step, we will begin by categorizing the data into types: text, images, or videos. Use tagging to add metadata, which helps in organizing the data and makes it easier to access and analyze later.

Then, and after that, particularly for text data, we'll use natural language processing (NLP). We will employ NLP techniques to extract useful information from text. Tools like NLTK, spaCy, or even more advanced models like BERT can help you perform tasks such as sentiment analysis, entity recognition, and topic modeling.

After that, we will do text indexing. We can use Elasticsearch or Apache Solr to index large volumes of text. These tools provide powerful search capabilities and can handle complex queries efficiently.

After that, we'll move to image data. To structure image data, we'll use image processing. We'll use libraries like OpenCV for basic image processing tasks such as filtering and transformations. For more advanced image analysis, consider deep learning models using frameworks like TensorFlow or PyTorch.

Then we'll do feature extraction. Apply techniques to extract features from images such as edges, textures, or key points, which can be used for further analysis or machine learning.

And then we'll come to video data. Here we'll do video processing. We'll use tools like FFmpeg, which can be used for basic video processing tasks such as format conversion or extracting frames. For analyzing video content, look at machine learning models that can classify or recognize activities in the video.

After this, we'll move to temporal analysis for videos. Temporal components are important; techniques like sequence modeling or recurrent neural networks (RNNs) can be useful to analyze sequences of frames for activities or events.

And then we'll move to data storage and management. Given the volume and complexity of unstructured data, use big data platforms like Hadoop or cloud services like AWS S3 for storage. These platforms can scale up to handle large data sizes and provide the necessary infrastructure to store and retrieve unstructured data efficiently.

And then we have visualization and reporting. We'll create custom dashboards. We will develop custom dashboards using tools like Tableau or Power BI, which can integrate different data types and provide a unified view of the analyzed data.

After that, we will do data summarization. Tools that provide summarization capabilities can help in condensing large volumes of unstructured data into more manageable and interpretable forms.

After that, we'll leverage these strategies and tools and can effectively manage, analyze, and derive insights from unstructured data, which can be crucial for making informed decisions in various applications. This is the path that you can explore and explain to the interviewer if this question has been asked.

Now moving to question number 10, and that will be based on scaling AI solutions in an enterprise. The question is: your company wants to scale its AI operations from a few initial pilot projects to enterprise-wide implementation. What are the key considerations and steps you would take to ensure the successful scaling of AI solutions across the organization? What challenges might you face, and how would you address them?

You can start answering this question with scaling AI solutions. You could answer that scaling AI solutions across an enterprise requires careful planning and strategic implementation to ensure success and alignment with business objectives. There should be a strategic approach to implement this. So starting with the approach, the number one step will be strategic alignment.

Identify business objectives. Start by identifying the business objectives that the AI solutions are intended to support. This ensures that the AI initiatives are aligned with the company's strategic goals and can demonstrate clear business value.

And then comes stakeholder engagement. Engage stakeholders from various departments early in the process to gather input and build support. This helps in understanding diverse needs and ensures broader acceptance of the AI solutions.

After that comes infrastructure and technology. Assess and upgrade infrastructure. Evaluate whether your current IT infrastructure can support the expanded use of AI. You might need to upgrade hardware, invest in cloud solutions, or adopt technologies that facilitate AI processing and data handling.

After that, we have standardization of tools. Standardize the tools and platforms used for AI development to ensure compatibility and ease of maintenance across the organization.

After that, we'll move to data management. Implement a robust data governance framework to manage enterprise data effectively. This includes policies for data quality, security, and compliance, especially important when scaling AI solutions that rely on vast amounts of data.

After that, we will come to data accessibility. Ensure that data is accessible across the organization but also secure against unauthorized access. This involves setting up secure data lakes or warehouses that centralize data while allowing controlled access.

And then we come to the next step, that is talent and training. Build AI competency. Develop in-house AI expertise through training programs and hiring. This builds the necessary skills within the organization to develop, manage, and scale AI solutions.

After that, you can also form cross-functional AI teams. Form cross-functional teams that include data scientists, IT professionals, and domain experts. This fosters collaboration and ensures that AI solutions are developed with a comprehensive understanding.

After forming these collaborative teams, we'll move to scalable deployment models. Pilot test and phase roll out. Before a full-scale rollout, conduct pilot tests to gauge the AI solution's effectiveness and integration capabilities. Based on feedback, adjust and then gradually deploy the solutions across the organization.

And then we have modular and flexible design. Design AI systems to be modular and scalable, allowing for adjustments and expansions as needed.

Then we'll monitor and do continuous improvement. Establish performance metrics to regularly assess the performance of AI systems. We will monitor these systems to ensure they meet expected outcomes and adapt as necessary.

After that, we have the next step, that is addressing challenges. There could be cultural resistance; there could be employees who resist the changes, but we have to address this through continuous education and by showcasing successful AI use cases within the organization.

By carefully considering these aspects and methodically implementing steps, you can successfully scale AI solutions across your enterprise, driving significant business value and innovation.

And that's all for question number 10. Now we'll move to question number 11, and that is based on ethical considerations in data science. The question is: in your data science projects, how do you ensure that ethical considerations are addressed? Describe the steps you take to identify and mitigate ethical risks in your projects. What frameworks or guidelines do you follow?

You could start answering this question with ethical considerations. They are crucial in data science to ensure that the solutions and analyses do not inadvertently cause harm or bias. Here's how you can ensure that. There are some steps, and we will discuss those steps.

Starting with number one, that is educate on ethical standards. Stay informed about the ethical standards in data science, such as fairness, accountability, transparency, and privacy. Organizations like the Data Science Association and the ACM have codes of ethics that we refer to as guidelines.

And then we have ethical risk assessment. Identify potential ethical issues. At the beginning of each project, conduct a thorough assessment to identify any potential ethical risks, such as biases in data or impact on vulnerable groups. This involves reviewing the source of data, the methodologies used for data collection, and the intended use of the data analytics results.

And then we have stakeholder analysis. Engage with stakeholders to understand the diverse perspectives and potential impact of the project. This helps in identifying ethical issues that may not be apparent from a purely technical standpoint.

And then we have mitigation strategies. Implement bias mitigation techniques. We will use statistical and machine learning techniques to detect and mitigate biases in data. This might involve techniques like resampling, re-weighting, or using algorithms designed to be fair.

And then we have privacy-preserving methods. Employ methods such as data anonymization, encryption, or differential privacy to protect individual privacy when analyzing sensitive data.

Then we have other methods: transparency and explainability. We have model explainability.

And after that, coming to documentation and reporting. Maintain thorough documentation of data sources, model decisions, and methodologies.

And then we have continuous monitoring and feedback. Monitor outcomes, and feedback mechanisms should be applied.

And then we have panels: collaboration and advisory panels. Then we have ethical review boards. For complex projects, setting up or consulting with an ethical review board can provide oversight and diverse perspectives on the ethical implications of project methodologies.

So by proactively addressing ethical considerations through these steps, you can ensure that your data science projects uphold high ethical standards and positively contribute to society while minimizing harm.

So this was all about question 11. Now moving to question number 12, that is based on time series forecasting for business decisions. The question is: you are tasked with forecasting monthly sales for a retail company using time series data from the past 5 years. What steps would you take to prepare and analyze this data to make accurate forecasts? What specific tools or techniques would you use and why?

We can start this by explaining time series forecasting. It's a powerful tool for predicting future events based on past data, especially in business contexts like retail sales. So we will have a structured approach here. We'll start with data collection and cleaning.

First, you will gather data. Ensure that you have collected all relevant data, including monthly sales figures from the past five years. Also consider including external factors that might affect sales, such as economic indicators, holidays, and promotional activities.

Then we'll proceed to clean data. Check for and handle any inconsistencies or missing values.

And then we have data visualization. We will plot the data. We'll use plotting libraries like Matplotlib or Seaborn in Python to visualize the data. This will help in identifying patterns, trends, and seasonality.

And then we have decomposition of data. There's seasonal decomposition. We'll use statistical techniques to decompose the data into trend, seasonality, and residuals. This can be accomplished with tools like the seasonal_decompose function from the statsmodels library in Python. We'll understand these components separately and can improve the accuracy of our forecast.

And then the next step is model selection and forecasting. There are two models: ARIMA and SARIMA models. Choose appropriate forecasting models based on the data's characteristics. For instance, ARIMA (autoregressive integrated moving average) is effective for non-seasonal data, while SARIMA (seasonal ARIMA) is suitable for data with seasonal patterns.

After choosing the model, we'll move to cross-validation. We will implement time series-specific cross-validation techniques like time-based splitting to evaluate model performance. This will ensure your model generalizes well on unseen data.

And then we have model fitting and diagnostics. We will fit the model. Using the SARIMAX class from statsmodels, we'll fit your model to the data. We will carefully select parameters based on AIC (Akaike information criterion) scores or thorough grid search techniques. Then we can do the diagnostics and forecast and validation.

After forecast validation, we'll move to iterative improvement. There's a feedback loop that should be mandatory. Regularly update the model with new sales data and refine the model as needed. This continuous improvement cycle helps adapt to changing patterns in sales data.

By following these steps and using these tools, you can create robust forecasts that help the retail company plan better and make informed decisions.

So this was all about question number 12. Now moving to question number 13, that is based on customer segmentation using machine learning. The question is: you are given a data set containing demographic and purchasing behavior data for a group of customers. Your task is to segment these customers into distinct groups based on similarities in their purchasing behavior and demographics. What steps would you take to perform this segmentation? Can you provide a sample Python code snippet to illustrate the initial stages of data handling and model application?

We can start this by explaining customer segmentation. It's a powerful approach to tailor marketing strategies and improve customer service by identifying distinct groups based on their behavior and characteristics. And here also we have a detailed approach for this task.

We'll start with the number one step, that would be data exploration and pre-processing. There will be initial exploration. Begin by examining the data set to understand the features available, such as age, income, purchase frequency, etc. Then we'll look for missing values or anomalies and decide how to handle them. That could be using imputation.

Then we'll move to feature engineering. We will create new features that might be useful for segmentation, such as customer lifetime value or average transaction amount. We'll also use normalization. Normalize the data to ensure that one feature doesn't disproportionately influence the model due to its scale. We'll use standard scaling or min-max scaling as appropriate.

So then we'll come to the next step, that is choosing the segmentation technique. Here we have K-means clustering. This is a popular method for customer segmentation. We will decide on the number of clusters by using techniques like the elbow method or silhouette analysis to determine the optimal cluster count.

And then we have model implementation. In that, we will use data preparation. We'll prepare the data by selecting the relevant features and applying any final transformations. Then we have model fitting. We fit the K-means clustering model to the data and evaluate and interpret analyzing clusters. After analyzing clusters, we'll move to the next step, that is strategic insights. We will provide actionable insights based on cluster characteristics, such as targeted marketing strategies for each segment.

And then we have iterative refinement: feedback incorporation. We'll use business feedback to refine the segmentation. If additional data becomes available, incorporate it to enhance the model.

We'll import the libraries and modules. As you can see on the screen, we have imported pandas, random forest classifier, train_test_split, standard scaler, classification_report, and after that we will load the data; and for that, we have used the pandas to read the data, that is read_csv. And after that, we are processing the data, that is data pre-processing. We are handling missing values and using the forward fill or fill to fill missing values in the data set. And then we are feature scaling, that is normalizing the selected features, that is feature one, feature two, and feature three, using standard scaler. And then we'll move to the next step, that is data splitting.

We'll split the data set into training and testing sets. So that test_size equal to 0.2 parameters specifies that 20% of the data will be used for testing. And then we'll train the model. We'll initialize and train a random forest classifier with 100 trees and a random state for reproducibility. And then we'll evaluate the model. We will make predictions on the test set, that is X_test, using the train model and print a classification report showing precision, recall, F1 score, and support for each class. So this code demonstrates the process of loading, pre-processing, training, and evaluating a machine learning model, that is random forest classifier, for predicting equipment failures in a manufacturing plant. The use of techniques such as data pre-processing and splitting along with the random forest classifier highlights a standard flow for building predictive maintenance models. So this was all about the question number 13.

So now move to the question number 14, that is based on predictive customer churn, and the question is: you are tasked with developing a model to predict which customers are likely to churn from a subscription service. So what steps would you take to build this model, and can you provide a sample Python code to illustrate the data preparation and model training process? So we'll start answering this question about depicting what is predicting customer churn. So predicting customer churn is crucial for businesses to implement detention strategies proactively, and we'll have a detailed approach for building a predictive model for this purpose. Starting with data collection and exploration, and in this we will collect data and after that we'll perform the exploratory data analysis, that is EDA. We'll perform an initial analysis to understand patterns and trends. And then we have feature engineering. We will create new features and derive new features that might influence churn, such as change in usage pattern or service upgrades. And then we'll handle the missing values if we found any. And then we'll encode categorical variables. We'll use techniques like one-hot encoding or label encoding for categorical variables. And then we have scale features to normalize or standardize numerical features to ensure they contribute equally to the model's performance. And then we'll select the model, that is we'll choose the appropriate model and start with for the knowing handling binary classification task that could be with logistic regression, random forest, or gradient boosting machines. And after selecting the model, we'll train the model and evaluate it. So fit your model on the training data and after that evaluate the model using appropriate metrics like accuracy, precision, recall, F1 score, and ROC to gauge its performance. And then we'll optimize the model using hyperparameter tuning. We'll optimize the model parameters using grid search or random search to improve performance. And then we have feature importance, that is analyze and rank features by their importance in predicting churn to refine the model further. And then and then the last step is deployment and monitoring. We'll deploy the model once validated, deploy the model into a production environment where it can predict real-time churn. So after deploying the model, regularly monitor the model to ensure it remains effective over time as new data comes in.

So now we'll see the sample Python code for this example. So starting with the importing of libraries, we will import pandas, numpy, scikit-learn, sklearn, tensorflow, and the tensorflow keras and callbacks. And after importing the modules, we'll start with data loading. We'll load the data set from a CSV file named equipment_data.csv, and that too with the pandas data frame. And after that, we'll do the data pre-processing. We'll handle missing values, and for that, we'll use forward fill to fill missing values val in the data set. And then we have feature scaling that will normalize the selected features, that is feature one, feature two, feature three, using standard scaler. And after that, we'll use the data splitting. We'll split the data set into training and testing sets, and the test_size will be equal to 0.2, 2, and this parameter specifies that 20% of the data will be used for testing. And after that, we'll start with building the model. First, we'll see sequential model that initializes a sequential model technique. And then we have dense layers that adds two dense layers with 64 units and ReLU activation function. Then we have dropout layers that adds two dropout layers with a dropout rate of 0.5 to reduce overfitting. After that, we'll do the model compilation. We'll compile the model using the Adam optimizer and binary cross entropy loss function for binary classification, and there will be an early stopping that will define an early stopping callback to stop training when the validation loss metric has stopped improving after three epochs. And after training the model, we will evaluate the model, evaluating the model on the test data and print the loss and accuracy metrics. So this code demonstrates the process of loading, pre-processing, building, compiling, training, and evaluating a deep learning model using tensorflow and keras for predicting equipment failures in a manufacturing plant. So the use of techniques such as data pre-processing, dropout regularization, and early stopping helps in building a robust deep learning model for predictive maintenance. So that's all with question number 14.

Now we'll start with question number 15, that is based on deep learning and NLP. And your question is: you are tasked with developing a sentiment analysis model using deep learning to understand customer opinions from reviews. So what steps would you take to build this model, and can you provide a sample Python code snippet to illustrate how you would pre-process data and train a simple deep learning model? So we'll start answering this with sentiment analysis that sentiment analysis using deep learning allows businesses to gauge customer sentiment from text data like reviews or comments effectively. And we'll have a detailed approach for building a sentiment analysis model. We'll start with data collection and cleaning. We will collect the data, gather a substantial data set of text reviews and their associated sentiments, typically labeled as positive, negative, or neutral. And then we'll clean the data, pre-process the data by removing noise such as HTML tags, special characters, and stop words. And we'll normalize the text by converting it to lower case. And then we have text pre-processing. We'll convert text into tokens, words, or phrases. And then we have vectorization that transforms tokens into numerical format using techniques like word embeddings or TF-IDF, that is term frequency inverse document frequency. And then we'll use the padding. And then we have the option of model selection. We'll choose a model architecture based on a basic approach and use an RNN or more advanced architecture like LSTM, that is long short-term memory, or GRU, that is gated recurrent units, which are effective for sequence data like text. And then we have model training. We'll compile the model, define the model architecture and compile it with a loss function suited for classification like categorical cross entropy and an optimizer like Adam. And then we'll train the model. We'll fit the model on our pre-processed data. We'll evaluate and optimize it. Evaluating model performance. Here use the metrics such as accuracy, precision, recall, and F1 score to assess the model. And then we have hyperparameter tuning. We'll optimize the model by adjusting parameters like learning rate, number of layers, and units per layer. And then coming to deployment, we'll deploy the model and integrate the model into the existing review processing pipeline. So it can automatically classify new reviews.

So let's see the sample Python code, and we'll have a basic approach for that. Here we'll import numpy, tensorflow, sequential, embedding, LSTM, dense, dropout. So embedding converts positive integers, that is indexes, into dense vectors of fixed size. And LSTM, that is long short-term memory layer, that is used for learning dependencies in sequence data. And then we have dense, that is a regularly densely connected NN layer. And then we will import pad_sequences. And after that, we have the data set and the sample text data representing customer reviews that will store in variable text. And then we have labels that has binary labels indicating sentiment, one for positive, zero for negative. And now we'll start with the pre-processing of data. Here we have declared that tokenizer. We will initialize a tokenizer that will handle only the top thousand most frequent words. And then we have fit_on_texts that is update the internal vocabulary based on the list of texts. It essentially creates a dictionary of word to index pairs. And then we have texts_to_sequences that will transform each text in text to a sequence of integers. And then we have pad_sequences that will ensure all sequences have the same length by padding shorter sequences with zeros up to the maximum length. And then we'll start building the model. Here we have sequential model that will set up a linear stack of layers. And then we have embedding layer that will map each word index to an embedding vector of size 64. So the input_length is set to 10, that is the length of the input sequences. Then we'll start with LSTM layers. So two LSTM layers are added. The first one returns sequences to allow the next LSTM layer to process these sequences. And after that, we have the dropout layer that applies dropout with a rate of 0.5 of the first LSTM layer to reduce overfitting. And after that, we'll come to dense layer that has output of a single scalar that represents the predicted sentiment and using sigmoid activation to output a probability. And now we'll start with model compilation and training. So we'll configure the model for training, and we'll use binary cross entropy as the loss function, that is suitable for binary classification, and the Adam optimizer and tracks constantly accuracy as a metric. And then we have the fit that trains the model for a specified number of epochs, that is iterations over the entire data set. And then we'll predict the model, that is after training the model can predict the sentiment of the reviews in the data set. This is useful for checking how the model performs on the training data itself. So this breakdown explains each step of the coding process, detailing how the data is prepared and how the model is configured, and then we'll compile it and use for training and prediction. So it's detailed explanation should help in understanding how to implement a simple LSTM model for sentiment analysis in TensorFlow. Now moving to the question number 16.

So let's start with question number 16, that is based on anomaly detection in transaction data. So the question is: you are tasked with identifying unusual transactions in a company's financial data that might suggest fraudulent activity. So what steps would you take to develop an anomaly detection model, and can you provide a sample Python code snippet to illustrate how you would pre-process the data and apply an anomaly detection technique? So we'll start answering this with anomaly detection technique, that is anomaly detection is essential for preventing fraud by identifying transactions that deviate significantly from typical patterns. And now we'll see the structured approach to building an anomaly detection model for transaction data. We'll start with data collection and cleaning, and we'll collect all the compiling transaction data, which should include details like transaction amount, time, user ID, and transaction type. Then we'll move to feature engineering and develop features that capture the essence of transaction such as time of day and the day of the week. And then we have data normalization. We'll use scaling techniques such as min-max scaling or standardization to ensure that the model is perfectly normalized. And then we have choosing the anomaly detection technique. So here we have to choose the technique which is effective for high-dimensional data sets and works for isolating anomalies instead of profiling normal data points. After choosing the anomaly technique, we'll train anomaly identification. We'll fit the chosen model to the data, and the anomalies that would have been chosen will be those transactions that the model identifies. And after this, we come to the last step, that is review and action. Here we have manual review, that is transactions flagged as potential anomalies should be reviewed manually to confirm fraudulent activity. And then we have continuous improvement, that is we can regularly update the model with the new data and feedback from the review process to improve accuracy. And now moving to the prediction, that is after training the model, we can predict the sentiment of the reviews in the data set, and this is useful for checking how the model performs on the training data itself.

Now we'll see the Python code to see how you can set up this model for anomaly detection. We'll start by importing the libraries and modules, and after that we'll load and prepare data. That is we'll load transaction data from a CSV file into the pandas data frame. And after that, we'll convert the transaction time column to datetime format, which allows the extraction of additional time-based features. And after that, we'll perform feature engineering that will extract the hour of the day from the transaction time column. This feature can be important as transactions occurring at unusual hours may be indicative of fraud. And then we'll move to the normalization of data. This will apply standard scaling to the amount and hour of the day feature. This normalization process involves subtracting the mean and dividing by the standard deviation for each feature, ensuring that the features contribute equally to the analysis and improving the performance of many machine learning algorithms. And after that, we'll start with anomaly detection with isolation forest. That's a technique. We'll initialize an isolation forest model with 100 trees, that is n_estimators equal to 100, setting the proportion of outliers, that is contamination, to 1% of the data. So this parameter is crucial as it influences the threshold of marking an observation as an anomaly. Then we fit the model to the scaled amount and hour of the data and predict the anomaly status for each transaction. And then we'll start with filter and display anomalies. We'll filter out transactions identified as anomalies, that is anomaly == -1. We'll display these transactions, which can be reviewed manually to determine if they represent actual fraudulent activity. So this code snippet provides a systematic approach to detecting anomalies in transaction data, leveraging the isolation forest algorithm's ability to handle complex and high-dimensional data sets effectively. So the pre-processing steps ensured that the data is appropriately formatted and normalized for optimal model performance. So this was all about question number 16.

Now moving to question number 17, and that is based on integrating machine learning models into web applications. And your question is: you have developed a machine learning model to predict real estate prices based on various features like location, size, and amenities. How would you integrate this model into a web application to allow users to get real-time price predictions? Can you provide a sample Python code snippet to illustrate how you would prepare the model for integration and handle user requests? So starting with the approach, that is integrating a machine learning model into a web application. This will involve several steps to ensure the model is accessible and performs well in a live environment. So here's how you can approach this task. We could divide into steps, and we'll start with number one step, that is model preparation. We'll finalize and save the model. So once your model is trained and validated, save it using a format that can be easily loaded into a web application. So Python's pickle module or TensorFlow's save_model format are commonly used for this purpose. Then we can use web application backend setup. For this, select a suitable web framework. So Flask is popularly known for its simplicity and effectiveness in integrating Python-based machine learning models. And after that, we'll develop the API. After developing the API within your Flask app, that you can receive user inputs for model features, load the model, make prediction, and return the result. And after this, we'll develop the UI. We'll design a user-friendly interface. We'll create a simple and intuitive UI that lets users input the features like location, size, and amenities and submit them for prediction. And after that, we'll move to the deployment phase. We'll use a cloud platform like Heroku, AWS, or Google Cloud to deploy your Flask application. And then we have the maintenance and updates. We'll monitor and update regularly for the performance and use the model as needed based on user feedback.

So now moving to the Python code and see how this model can be created. So here we'll start importing the libraries and modules, and we are using flask, pickle, and jsonify, and we will start with app initialization. We'll initialize a new Flask web application. That would be a special variable which gives Python files a unique name to differentiate between them when they are imported into other scripts. And after that, we'll load the model. So loading a pre-trained machine learning model from the file system. So this model is assumed to be saved in the same directory as this script. So the model is loaded in 'rb' mode, which stands for read binary. And after that, we'll move to API route and prediction function. So we will define an API endpoint at '/predict' that listens for POST requests. This is the URL that the front end of the web application will call to send data to the back end. And after that, we'll start with predicting the function. And here we have extract_features that retrieves data sent in JSON format from the POST request, that is request.get_json(), and the force we have set it as true here and forcefully formats the request data into JSON, ensuring compatibility. And then we'll extract the relevant features, that is location, size, and amenities, from the JSON object and store them in a list as expected by the model. And after preparing the features, we'll make the prediction. We'll use the loaded model to make a prediction based on the provided features. And then we have the return_prediction method. Here we will convert the prediction result into JSON format using jsonify and send it back to the client. And this will ensure that the response can be easily handled by the client application. So this was all about the question number 17.

Now moving to the question number 18, that is based on analyzing. And now we'll move to the question number 18, that is based on analyzing geospatial data. And your question is: you are tasked with analyzing geospatial data to help a city improve its public transportation system. The data includes GPS coordinates of bus stops, ridership numbers, and traffic patterns. What steps would you take to analyze this data, and can you provide a sample Python code snippet to illustrate how you might visualize bus stop location and ridership? So you can start answering this question that geospatial data analysis can provide critical insights into how effectively a public transportation system serves its city and guide improvements, and there's a detailed approach for this, and we can start with data preparation, and in this we'll do data collection and data cleaning. And after this step, we'll move to the next step, that is exploratory data analysis, and in this we'll have statistical summary, we'll generate descriptive statistics, and then we have correlation analysis. And after moving that, we have geospatial visualization, that is mapping bus stops. We'll plot the locations of bus stops on a map to visually assess their distribution across the city. And after that, we have heat maps that will create ridership data to identify hotspots and areas with potential service gaps. And after geospatial visualization, we'll move with spatial analysis. We have proximity analysis that will analyze the proximity of bus stops to key areas like commercial centers or residential areas. And now moving to the fifth step, that is optimization and recommendation. So we'll have route optimization that will suggest modifications to routes based on traffic patterns and ridership demand, and the policy recommendations that will provide actionable recommendations for improving bus frequencies. Now move to the sample Python code where we can define this model and use it accordingly. And here we will start importing the libraries and modules. And here we'll start with importing geopandas and matplotlib.pyplot. And after importing, we'll start with data loading. So we will declare a variable bus_stops and load the bus stops data from a shapefile. So shapefiles are popular geospatial vector data formats for geographic information system software. And then we have the ridership that will load ridership data from a CSV file, which includes columns for longitude, latitude, and ridership levels. And after that, we'll create geo data frame that will convert the ridership data frame into a geo data frame, and this step involves creating a geometry column from the longitude and latitude columns. And then we have the plotting one here. We will plot the graphs that would be figures and axes, and create a figure for the single subplot with a specified size, that is 10x10 inches. And then we have city_map.plot. It is assumed that there is a base map of the city loaded as a geo data frame named city_map. This is plotted first with a light gray color to serve as a background for the other layers. So this was all about question number 18. Now moving to the question

Number 19: That is based on predictive maintenance using machine learning. Your question is: You are tasked with developing a predictive maintenance system for a manufacturing plant that relies heavily on automated machinery. The data available includes machine operational parameters, maintenance history, and failure incidents. What steps would you take to develop a predictive model, and can you provide a sample Python code?

So you can start with predictive maintenance, which is essential in manufacturing as it helps prevent equipment failures, reducing downtime and maintenance costs. And here you would have a detailed approach or predictive model for this, starting with data collection and integration. Then you can do EDA (exploratory data analysis), and then we can perform feature engineering, and then move to data pre-processing tasks, and then the model selection and training. After that, we have model evaluation and deployment techniques that we can do for the model using appropriate metrics such as precision, recall, and F1 score.

So this was all about question number 19.

Number 20: That is based on personalization using machine learning. Your question is: You are tasked with developing a machine learning model to personalize content recommendations for users on a media streaming platform. The data available includes user demographic details, viewing history, and ratings. What steps would you take to build a model for personalized recommendations, and can you provide a sample Python code for that?

So you can start answering this with creating a personalized recommendation system. This would be essential for engaging users by providing content that is relevant to their interest. And there will be a systematic approach for personalized content recommendation. We'll start with data collection and integration. And after that, we'll perform EDA (exploratory data analysis), and then we have feature engineering. In this, we'll interact features and the temporal features. We'll include time-based features to capture trends and seasonality in viewing behavior. And then we'll select the model, that is, by collaborative filtering and hybrid models. And then we'll train the model and validation and implement and monitor them. And after that, we'll deploy the model and integrate the recommendation system.

So that's a wrap on this full course, guys. If you have any doubts or questions, ask them in the comment section below. Our team of experts will reply to you as soon as possible. Thank you and keep learning with Simply Learn.

Staying ahead in your career requires continuous learning and upskilling. Whether you're a student aiming to learn today's top skills or a working professional looking to advance your career, we've got you covered. Explore our impressive catalog of certification programs in cutting-edge domains, including data science, cloud computing, cyber security, AI, machine learning, or digital marketing. Designed in collaboration with leading universities and top corporations and delivered by industry experts. Choose any of our programs and set yourself on the path to career success. Click the link in the description to know more.

Hi there. If you like this video, subscribe to the SimplyLearn YouTube channel and click here to watch similar videos. To nerd up and get certified, click here.