📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Linear Algebra for Machine Learning

freeCodeCamp.org10:48:43

Transcription

This in-depth course provides a comprehensive exploration of all critical linear algebra concepts necessary for machine learning. You'll learn the mathematical foundations to excel in AI. Tdiv from Lunar Tech developed this course; she has created many popular machine learning courses.

Machine learning is at the forefront of the innovation powering the most advanced and transformative systems for companies like Apple, Tesla, Netflix, Amazon, OpenAI, and many others. It enables the creation of intelligent systems that can predict trends, personalize user experience, and automate complex tasks. To develop these practical applications, a deep understanding of the underlying mechanics is important. This requires a solid grasp of the mathematics behind machine learning, so all these technical details with a particular focus on linear algebra. This all-encompassing course explores linear algebra in an interactive and machine learning-focused manner.

Welcome to the Linear Algebra for Machine Learning course. You will acquire the critical principles needed to build, optimize, and analyze sophisticated machine learning models, from designing customer algorithms to enhancing current technologies. This course provides the mathematical foundations with vital interest for those pioneering advancements in machine learning, for those dedicated to mastering the mathematical aspect and the technical details behind machine learning. Our extensive 26-plus-hour course on fundamentals of machine learning within the mathematics boot camp, as well as a separate course, offers an in-depth exploration. This extensive program includes certification and is tailored for individuals serious about advancing their career in the field of machine learning and engineering. This crash course in mathematics will serve you as a great starting point by establishing a robust foundation in linear algebra. You will be well prepared to excel as a machine learning practitioner, equipped with the mathematical knowledge that drives the innovation and efficiency in this field. So if you're ready, I'm really excited, and without further ado, let's get started.

Welcome to the course on the fundamentals of linear algebra presented by Lun Tech Academy. My name is D Vasan, and today we are going to start with some basic concepts important for understanding linear algebra. Linear algebra is one of the most applicable areas of mathematics. It is used by pure mathematicians that you will see in universities doing research, publishing research papers, but also by the mathematically trained scientists of all disciplines. This is really one of those areas in mathematics that you will see time and time again appearing in your professional life if you want to become a job-ready data scientist or you want to do some hands-on machine learning, deep learning, and AI stuff. But also, linear algebra is used in cryptology; it is used in cybersecurity and in many other areas of computer science and artificial intelligence.

So if you want to become this well-rounded professional, you want to go beyond using libraries, and you want to truly understand the mathematics and the technical side of these different machine learning algorithms—from very basic ones like linear regression to most complex ones coming from deep learning, like architectures in neural networks, how the optimization algorithms work, how the gradient descent works, and all these other different methods and models—then you are in the right place because you must know linear algebra such that you will understand these different concepts, from very basic ones to most advanced ones in data science, machine learning, deep learning, artificial intelligence, data analytics, but also in many other applied science disciplines.

So before starting this comprehensive course that will give you everything that you need to know about linear algebra, first, I'm going to tell you what we assume that you already know because linear algebra, it comes from about the third year of Bachelor's of different highly technical studies, and here we are assuming that you already know certain concepts. So to ensure that this course stays really on the topic of linear algebra and that you understand all these concepts really well, for that we need to be able to know different topics. So before we dive into these concepts, let's familiarize ourselves with the basic prerequisites and notation used throughout this course, and you will really need to know this in order to understand these concepts really well, such that instead of memorizing, you'll actually just hear me once or maybe twice, and then every time you hear later on or you see it in the papers or in some algorithms, you will recognize this is something that we already learned.

So some key prerequisites overview is here. First of all, to fully grasp the upcoming material, you should be familiar with some basic concepts like real numbers, vector spaces. So you don't need to know this idea of vectors, though you already most likely are familiar with this given that you know how to plot different lines, you know the idea of x's and y's and how to plot these different graphs. But here we are going to touch base on this every time when we come close to these concepts. I will refresh your memory, and we will go through these numbers, the idea of norms and distance measures. Because when it comes to the vectors, when it comes to the magnitude and all these different topics that we are going to discuss as part of linear algebra, knowing what a norm is and what is the definition of distance, what is the length between two points when we plot it into two-dimensional space or three-dimensional space, those are all very basic concepts that usually use as part of a basic pre-algebra or just common algebra courses and lessons.

In order to truly understand what linear algebra is about, to understand the direction of vectors, the angle, and then the dimensionality reduction, how linear algebra is applied, for instance, in different algorithms in machine learning, deep learning, data science, statistics, you really need to understand this Cartesian coordinate system. So this is not only important for linear algebra, but I assume you already know it given that you have passed those other courses like calculus; usually they are covered as part of pre-algebra or algebra. So the Cartesian coordinate system—I mean here, understanding what is, for instance, the common description of them, for instance, when you when we write like X and then Y on the vertical axis, and then we can we have here zero, and then we can always plot these different plots; you know, we have a clear understanding what this Y is equal to X line, we understand how by knowing certain points we can plot different plots; for instance, that this is the Y is equal to X line, that here it means that if we have here one, then this is just one, two, this is two. So we understand when we have the function of the line and we have a certain value at is our y-coordinate or x-coordinate, then the corresponding coordinate can be found.

Then you also need to know some basic things that I just didn't mention right now. So, for instance, that the numbers here can be like 1, 2, 3 up to infinity; so you understand this concept of infinity, and then here the same story; then here we have minus one, you know, minus two, uh, and then this is then used later on, and we will be touching base on this one; we will be describing our vectors and how we can visualize our vectors either in two-dimensional space, like we have here, because this is two-dimensional, so we have X and Y, but we can also, of course, visualize it in three-dimensional, etc. So this idea of a basic coordinate system is really important, usually covered as part of algebra; if not pre-algebra. Then we have basic trigonometry, which means that you need to have a clear understanding what sine is, what cosine is, what tangent is, and their reciprocals. And here I mean that you know, for instance, what is a cosine function, what is a sine function; you know that you have an understanding, for instance, that what is this line, you know, whether it's a sine line or cosine line; you have also an understanding what this Pi is. One thing that I didn't mention, but it it just goes around all these topics, some basic things that you understand what is X, what is Y, why we use them, and this idea of variables. And also you need to understand this idea of a square or, you know, 90-degree angle, and then the Pythagorean Theorem. Here we have the same; so what is this relationship between different sides of a triangle that is a very unique triangle and that has one of the angles as 90°, and this idea of, you know, the sides, how this relates to the sine, cosine, tangent, cotangent, and also how the Pythagorean theorem applies when we have a triangle but it is is no longer with an angle that is 90°; what is the sum of all the angles of a triangle? So those are basic stuff that are commonly covered as part of trigonometric lessons or part of general geometry.

Then another prerequisite is this understanding of identities and equations in trigonometric lessons, something part of which I already covered, and this is goes around having a basic understanding of algebra and geometry; those are super important to understand more advanced techniques from linear algebra. Then we have, finally, this idea of orthogonality, perpendicularity in vectors. So this also comes from geometry and from trigonometric lessons; so you understand that if we have, for instance, the two lines that don't have any intersections, then we are talking about two orthogonal lines, and otherwise, for instance, if we have and the two lines like this, then we are talking about perpendicular vectors when you have two lines that are actually parallel, so they don't have any intersection and you won't find any point that is common for the two. So when it comes to this R, so as part of real numbers and vector spaces, R represents the set of all real numbers. So you can be dealing with, for instance, integers like 1, 2, 3; this can also this will also cover all the negative numbers like -1, -2, -3, but also the floating numbers like 1.223 and all the other numbers that you can think of; those are the set of all real numbers. So this is in one-dimensional space, right? So you can see that I'm writing just one number, you know, two, three, and other numeric numbers.

Then we have the idea of R2, R3, up to RN, where now all these numbers they represent represent in this case the N, it represents the N-dimensional Euclidean space. So when it comes to this idea of N-dimensional numbers, so for instance, R2 here we just mean 2D plane. So I'm pretty sure you are familiar with this idea of, for instance, x-axis and y-axis. Here we are dealing with a two-dimensional plane, so for every point that we can find here, we can describe them by assigning them a value X, so coordinate X and a coordinate Y; that's exactly what we mean by saying that the number can be represented in a 2D plane. So here we are dealing with this two-dimensional space; this is our two-dimensional Euclidean space, and every number in here that is part of this R2 can be pictured here, can be represented in this visualization. So, for instance, if I have this number and let's assume that the value on the x-axis is two and we can see here that the corresponding Y is zero, I can describe this number, which I will call A, I can describe this by writing down first the x-coordinate, which is two, and then the y-coordinate, which is zero. So I'm then saying that A, which is a point with x-coordinate 2 and y-coordinate 0, it is part of my R2 and it's part of my two-dimensional Euclidean space. When it comes to R3, a similar thing we can do with that, only in that case we need not just x-axis and y-axis, but we need to add our third dimension. So here, for instance, when it comes to the R3, then we need to do y-axis, we need to have x-axis, but also we need to have some z-axis, so such that every time, every point in the space we can then describe by x, y, and z coordinates. So if we write it in terms of the vector, something that we will see very soon as part of our first unit of this course, we will then need to represent every number in this three-dimensional Euclidean space by writing down first the x-coordinate, let's say one, and then y-coordinate, let's say another one, and then z-coordinate which is one, or even better, even easier, let's use 0, 0, 0, which means that we are dealing with this initial number, which is the center of this three-dimensional Euclidean space.

When it comes to the N-dimensional or the higher-dimensional spaces, it's much harder to visualize; therefore, usually when it comes to visualizations, we do usually we usually only visualize the one-dimensional, two-dimensional, and three-dimensional spaces; above then it just no longer does make sense to visualize it, but we definitely deal with them, and they are part of applied linear algebra. So understanding these spaces is very important for analyzing vectors for their interactions, and this holds not just for this two-dimensional and three-dimensional but really for multi-dimensional spaces. Let's now quickly define this idea of norm. So the norm of a vector, denoted by this V, which you can see kind of like similar to the absolute value from pre-algebra, you can see here that we have this double straight lines like from absolute value, then we have the name of the vector or the variable name that we are assigning to our vector, and then you might notice here on the top of this this arrow; this basically says that we are dealing not with just a variable but really we are dealing with a vector. This is really important because you can see that there makes a huge difference if we have, for instance, just V or V1, I have to say, or just V; those are really important things that you need to keep in mind when it comes to linear algebra and trying to differentiate vectors from a point. You will notice that when it comes to norm, we can represent it either by this notation or this; usually it's a common notation in machine learning or in data science with this two bars, and when we do this we automatically also know that we are dealing with a Euclidean distance; we call it also L2 norm, and this is something very common and usually used as part of regression, which is an application of linear algebra and it's used in regularization; so we are regularizing our machine learning algorithms. So when you get into machine learning, you will see time and time again this notation; so next time when you see this, then you know automatically that you are dealing with L2 norm, and L2 norm, which is also used a lot in machine learning, it is referring to the usage of L2 norm to uh in the regression and regression or L2 regularization is a very popular regularization technique as part of machine learning. So right now even you can see this intersection or linear algebra or this idea of norms in machine learning. All right, so now let's see why we call it actually L2 norm or often referred to as Euclidean distance. So Euclidean distance, you can see here, which is also in this case this V which describes the norm of the vector v is equal to square root, and then we have all these coordinates assuming that the vector comes from an N-dimensional space. So you can see here the RN, the V vector, the Euclidean distance or the norm of this vector v is equal to square root, and then V1 squared plus V2 squared plus and all this in between numbers plus VN squared. So here basically it means take square root of V1 squared, V2 squared, plus plus V3 squared, blah blah blah, plus VN squared. So basically take all the units that form this vector and then so are on this vector and use them, square them, and then add them, and then take the square root of that; that's the distance or I have to say the norm of this vector. So why this is important, this idea of norms and Euclidean distance, beside of being used in machine learning and why is it used? So norms, they provide a way to measure the size or the length of a vector in vector spaces, which means that when we want to measure a distance, a similarity relationship between, for instance, vectors, then it becomes much easier to use this idea. And Euclidean distance is not only used in regularization techniques like L2 regularization or regression, but it's also used in other machine learning or deep learning algorithms as a way to measure the distance or the relationship or the similarity between two different entities; those can be variables, those can be two people that we want to compare in our algorithm, or two entities. Um, for instance, the norms or the Euclidean distance, they are also used as part of the K-means algorithm, something that you might have heard, and if you follow later on the machine learning and the clustering section of machine learning, you will see that Euclidean distance is used as part of K-means algorithm that aims to cluster observations into different groups. So this also yet another highly applicable topic that you must know in order to understand different linear algebra topics but also machine learning topics.

Let's now talk about a simple topic that we must know about and refresh our memory very quickly before moving forward to our next topic that is a prerequisite for this course. So the Cartesian coordinate system is just a fancy word of describing this idea of X and Y or XYZ when we just want to visualize them and showcase these numbers related to the space. So we just learned and I just quickly was talking about this idea of X and and Y and how we can visualize that in a plane. So the Cartesian coordinate system is a framework for specifying points in a plane or a space using an ordered list of numbers. So we know, for instance, when we plot this, then here we need to put X and Y in our two-dimensional space R2, and we know that here in the middle we have zero and here we have 1, 2, 3, 4 and the same here 1, 2, and then 3, 4, which means that everyone that is in the industry, whether it's in mathematics, in physics, in data science or ML or AI, we all universally agree on this system; we know this is this ordered list of numbers, and we know that if we have, for instance, a point here, then for this point we know that the x-axis and y-axis is definitely positive even if we don't know the corresponding numbers. And then once we have more general lines here, so not general but specific lines, then we even know the exact coordinates and values here, and we definitely know that this number should be, so the x-coordinate should be between two and three, so first we have the two and then three and not the other way around. So this ordered nature helps us to understand how we can put all these different numbers and organize them in our two-dimensional space, and we also know the corresponding Y; so we know that, for instance, our Y is not minus three because it's lying in here in this part of our coordinate system and not somewhere here where the y-axis are negative. And why do we know that? Because it's an ordered list of numbers that we can visualize in this 2D plane. And here you also need to keep in mind and we need to remind ourselves about this idea of these four different parts that we got; so we have our here the first part, the second part, the third part, and then the fourth part of our coordinate system, and here we we are dealing with a two-dimensional plane, but if we were to deal with the three-dimensional plane, we no longer have just x-axis and y-axis where x-axis were on the horizontal and y-axis on the vertical, but we have our third line which is the Z; so we have now three different dimensions, so X, Y, and Z, and we are basically extending our two-dimensional plane to three-dimensional. So this system is fundamental for visualizing and working with vectors geometrically, so then we can just use this two-dimensional plane in order to visualize this vector, for instance, knowing what are all these points that appear on this vector, what is its direction, where is it headed, you know, what is the beginning, and then we can also find out all the, so the relationship of these vectors with all the other vectors; for instance, if we have another vector here, then we can use the coordinates of them and information about vectors to understand that we are dealing with two parallel vectors that don't have anything in common, so no intersection points, whereas to say if we have another vector like this, and we know that here we are dealing with perpendicular, you know, orthogonal vectors. So this is why this Cartesian coordinate system is important, and it's not just important for linear algebra, but just in general for mathematics and for data science and for AI, and you will see this coordinate system time and time again in different visualizations, even when you want to visualize the mean of your data or you want to visualize the probability distribution function describing your population from statistics or from data.

Science you want to visualize, for instance, how your optimization is working, or you want to visualize how your model is performing in terms of its evaluation matrix. For all these cases and for any visualizations, this idea of the Cartesian coordinate system is going to become very handy.

Let's now talk about this idea of angles and the idea of circles, radians, the pi, as well as this degree sign. This comes usually from geometry or trigonometry, and this is very important when it comes to vectors. Because when we have two different vectors, then we want to understand their relationship: do they form this less than 90°? Or are we dealing with a sharp corner, a sharp angle? Or with a 90° angle? So we are dealing with this type of vectors where we have, you know, 90°, or we are dealing with um this type of vectors when the angle is 180°, which is, by the way, uh something that we are referring to as Pi.

And here is one thing that is important: it's not just Pi, but it's Pi radians. Why? Because in mathematics we also have this idea of Pi, which is usually a number that is 3.14. So we should not confuse this Pi with pi radians. The relationship between the two is something that we have also seen as part of our pre-algebra and algebra courses. So if it's something that you want to just refresh your memory on, this will be super helpful to check our very initial course on um all these Basics, so pre-algebra. So this number comes from pre-algebra, and then this idea of Pi radians, and just in general all this information about what is 180°, what is this angle, what is 360°, and all the information that comes from trigonometry and geometry can be found in our corresponding course.

The next topic is the unit circle. The unit circle is highly related to this idea of radians, degrees, cosine, sine, but also understanding the Cartesian coordinate system will help you to understand the unit circle. So this also comes from trigonometry and geometry, and it's basically a fancy way of saying we have an x-axis, we have a y-axis, we have here zero. So our common Cartesian coordinate system, only we are trying to focus on this part of the system where we have here one, we have here one. So on the x-axis we have one, and then here minus one, here minus one for the y-axis, and here y, the Y is equal to 1. So we have here all these points, and then we have the circle with the radius of one. So here is this, you know, this is the radius, and here we plot this circle, and this will help us to understand these concepts of sinus, cosinus. You know, the Theta is just a variable that we use to describe the angle. And for instance, here we are dealing with 45°; this angle is 90°; this entire thing is 360°; and half of it, so this part only is 180°. So those are all important parts of understanding this idea of the unit circle.

So you might have already guessed that the unit circle refers to this idea that we have here one unit, here one unit, one unit, one unit, forming this entire circle, so with the radius that is equal to one. All right. So this is something that is very easy, and this comes from geometry and pre- and trigonometry. Uh, you also need to understand this concept of the sinus and cosinus, and how sinus and cosinus are related to this. What do we refer to by the sinus and cosine? You know, what is this? What are these points? So, for instance, we understand that here the x is equal to one and Y is equal to zero. So here this point is simply 1, 0. So this point, and then we have 2 Pi radians. So what is this idea of Pi? So we know that Pi radians is simply the 180°, which means that you also need to understand this concept of Pi/2, which is simply the 90°. So you can see here one thing that I forgot to mention: you need to understand this concept, the relationship between the Pi and so Pi radians and radians and this unit circle. You need to know that here the Pi divided by two is simply this angle, and then the entire Pi is this angle, and then this entire thing, the entire angle with 360°, is equal to 2 Pi. So 2 Pi radians is simply this entire thing. So those are very easy concepts that come from geometry and trigonometry, and if you want to refresh them, then head towards those courses, because this will help you to understand all this concept from scratch.

Let's now continue our refreshment when it comes to trigonometric identities. We just spoke about this unit circle, we talked about the sinus, cosinus. It's really important to relate this back to a bit more advanced topics coming from the same domain and from the same area of mathematics. And here we, we need to know this concept before learning linear algebra. A few other things that um would be really great if you know, but it's actually not a must to understand all these different topics, is the idea of the Pythagorean identity. So don't confuse this with the Pythagorean Theorem. This is the Pythagorean identity. So this one, that the square of the sine of an angle plus the cosine squared is equal to one, and all these different rules that go around the sine and cosine, and also the what is, for instance, the sine 2 Theta, which is equal to 2 sine of theta and cosine of theta. You know, those are all different rules that would be handy to know. And if you are so far, I assume that you also know geometry and fundamentals of trigonometry, which means that you also know these truths, but this might be just a great time to go ahead and quickly refresh your memory on these concepts, because those might become handy in your applied linear algebra and applied mathematics journey. But for now, I would say this is not one of the most important things to know to learn this and to go through this course, but just something to keep in mind.

So when it comes to the trigonometric equations, uh this can become very handy later on when we want to prove something in linear algebra. So to follow along, it's actually a good idea to know, for instance, what is, how you can solve these different equations. And this will go back and refer to the unit circle that we just saw. For instance, if the sine Theta is equal to 1/2, then you will need to quickly remember what is that angle for which the sine is equal to 1/2. Then you realize that is actually the angle where you take the Pi, and remember that Pi is equal to 180°, and that is the one corresponding to, and then Pi / 6 is simply 180 / 6, so this is basically the 30°. So those are things that you can do when you know, for instance, all these different sine and cosine, so you have memorized for these different angles. So what is the sine and cosine for 30°, for 60°? Um, let me actually remove this to make it easier. So this type of problem is very easy to solve when we keep in mind and we memorize what are these different values for sine and cosine when it comes to different angles. For instance, for the angle equal to zero, let me actually remove this and clean this part for better understanding. So if we have, for instance, 0 degrees, then we know that the sine for this is zero and the cosine of this is one. So we are basically dealing, so if I plot a unit circle, we are dealing with this number. So remember that sine and cosine, those refer to the Y and X on our unit circle. So keep this one in mind. So if the cosine Theta is then equal to 1 and the sine, so Y is equal to 0, we are dealing automatically with this number, with this number, and you can see that here the angle is also zero. So here we are dealing with 1 and 0 coordinate. So this is our cosine of zero angle, and this is then our sine of zero angle. So we automatically, even from this graph, can see very easily that the sine of 0° is equal to 0 and the cosine is equal to 1. All right.

So let's quickly also refresh our memory on a few other degrees. So for the 30°, which is simply the Pi / 6, so this is 30°, then the sine, or the Y-axis, is equal to 1/2, and the cosine, or the X, x value, x coordinate, is equal to the square root of 3/2. So we are dealing with this, this corner or angle, so 30°. So even from here you can see that the coordinates make sense, make sense. Then we have the Pi/4, for another famous value, which is corresponding to the 45°. It's simply this angle, and for this angle the x-axis, which is the cosine, so this number is equal to 1 / the square root of 2, and then for the sine, the so the y coordinate is equal to 1 / the square root of 2. As you might have guessed, because in this number the x-axis and y-axis is equal to is the same. So you can see that this distance and this distance is the same because we are dealing with this type of figure. So here we have 45°, here we have 45°. So this values are the same, and this is something that you would know, knowing the Pythagorean theorem. So then you can go ahead and refresh your memory for the 60°, so here I'm referring to the Pi/3, and then the 90°, which is the very easy case. This one obviously the x-axis is equal to zero, so here you should have zero, and the y-axis is equal to one, so here you should have one, and so on.

All right. So we went into quite a detail here, but I think this is a very important topic. Knowing this idea of trigonometric equations, identities, this idea of the unit circle are super important because they are highly applicable to different fields in artificial intelligence, data science, machine learning, and will definitely set you apart.

All right. Let's now talk about the law of sines and cosines. Those are things that I won't be going into too much detail. I just wanted to quickly showcase to you: if you want to get the proof of those, definitely check out our corresponding courses. But for here, I'm assuming that you already know. So you know the law of sines, which means that if you have this triangle, you know you have these different sides, so you have an angle A, the corresponding side is a, and then you have angle B, corresponding side is B, and then here C and the corresponding side is C, then you know that A / the sine of that angle is equal to B / the sine of that angle, and then is equal to C divided by the sine of that angle. So basically take this value, divide it by the sine of this angle, you know, right in front of it, is equal to taking this value and then dividing into the sine of this angle. So the proof of this law is outside, so out of the scope of this course, but knowing this will help you to understand different concepts. And then the law of cosines is simply saying take the side of a target angle. So in our triangle we have here a, we have here angle B and the C, and if we go and look into this specific angle, so angle C, just randomly picking one of the three angles, then the side right in front of that angle, so the C, C squared is equal to if we take this, you know, the other two sides forming that angle, so A and B, is equal to A squared, so this is just a constant, a distance of this side A, squared + B squared, so this side squared - 2 * A * B times the cosine of that angle. This is what we are referring to as the law of cosines. Quite easy. We are not going to prove it again. If you want to get the proofs, make sure to check our other courses on geometry and trigonometry.

We're almost done with the prerequisites. Just a quick refreshment. We saw already the norm. Here is just an example what a norm is, and on a specific two-dimensional vector, when we have, for instance, that a vector is equal to three and four, which means for the first dimension, let's say on the x-axis we have three, and then on the y-axis is equal to four, then the norm, or the Euclidean distance, so this is equal to we take the x value, so three, and then we square it. So V, you can see here this is the case when n is equal to 2. This is simply equal to the square root of V1 squared + V2 squared. And as V1 is equal to 3, so this is our, maybe I can make this just V1 and this is my V2, then the norm or the Euclidean distance for this vector, so this thing is equal to V1 squared + V2 squared, which is equal to 3 squared + 4 squared, and this value is the square root of 25 and it's equal to 5.

Let's now see the difference between Euclidean distance and the norm. You could see here the norm, here we have just one vector, like here, and this norm it has just two corresponding values into two-dimensional space. You see here we have just three and then four, so this is V1 and V2. When it comes to the Euclidean distance, this is kind of the generalization of this idea of norm. So the Euclidean distance between two points A and B in R<sup>n</sup>, so in the N-dimensional space, is the norm of the vector connecting A to B. So we see that the norm and the Euclidean distance are highly related to each other. Only we are talking about the norm when it comes to one vector, but when we have this vector A and the vector B, this is simply the Euclidean distance. So for the Euclidean distance we know already this idea of distance, how we can measure it, and you can see that this comes very similar to what we see here, notation. And here we are saying, well we have this vector and then it has the two coordinates in n is equal to 2, in two-dimensional space. When it comes to the Euclidean distance, Euclidean distance helps you understand what is this distance between two points in an N-dimensional space. So the Euclidean distance between two points, let's say A and B in N-dimensional space is the norm of the vector connecting A to B. So for instance, if we have a point A and we have a point B, we are connecting this, and this is the vector connecting these two points, then the Euclidean distance is simply the norm of this vector. So this is the Euclidean distance. So we can see that norm and the distance, they are highly related to each other. In the Euclidean distance we are using this idea of norm, and specifically the norm 2, as I mentioned before.

So here you can see that the definition of Euclidean distance, so the distance between A and B, the two points, is equal to the square root of A1 - B1 squared + A, and then here we have basically A2 - B2 squared, and then plus A3 - B3 squared. Those are things that we cover as part of this dot dot dot, and then plus up to the last point when we have An - Bn squared. So here what we mean basically is that if we have two points, here is A and here's B, and this is vector, and we know all these different points, so A1, B1, A2, B2, A3, B3, blah blah blah, and then here An, Bn, we know all these points lying here in this distance, then we are taking them and using them to calculate the Euclidean distance. So here, for instance, if we have point A and B, so in this example, let's do a quick one specific example when we have a point A which has coordinates 1 and 2, so this is basically A1, A2, and then point B with points in it like B1, B2. You can notice that the d<sub>AB</sub>, so the distance or the Euclidean distance of these two points, which is equal to the norm of this vector, or here this is A and this is B and this is this vector, this is equal to the square root of 4 - 1, so it takes the B1, so this is B1 and this is A1, takes the square, and then says plus B2 - A2 squared, takes the square root of that, and says this is equal to 5.

Now you might be wondering, but hey, why do we do then, instead of 1 - B1 squared, we do B1 - A1 squared? And the answer to this question lies in the uh properties that we learn as part of pre-algebra, because it doesn't matter when we take A1 - B1 squared or B1 - A1 squared, because this squared ensures that it doesn't matter which one we take first and subtract the other. Now the proof of that is outside of the scope of this course, this is part of pre-algebra, but I just wanted to put this out there to ensure that you are seeing what we are seeing here, because here it says A1 minus B1, but in this example we are taking instead that B1 and we are subtracting A1. This is a common thing that we do in pre-algebra and just in general in different computing distance or distance-related cases. So I just wanted to put this here to ensure that later on this is something that can be clear from the first view, right? And in here we will quickly refresh our memory on the Pythagorean theorem, which basically says in the right angle triangle, so if we have this type of triangle, so here we have 90°, this is a right angle triangle, the square of the length of the side opposite to the right angle, so this side, this we often refer to as C and this as B uh and then A, those two are not very important, but this is commonly referred to by C, so the side opposite to the right angle, then we know that the square of the C, so C squared is equal to A squared + B squared. This is a super important theorem and a fundamental principle for defining the norms, the distances in Euclidean spaces, and in many other applications.

So the angles play a crucial role in understanding the direction of the vectors and you know how they can be measured in degrees or in radians. We saw also the Pi radian, this idea of, you know, that the Pi radian is equal to 180°. Those are all very important when it comes to linear algebra and just in general application of mathematics in machine learning, in AI and other applications. The relationships between these angle measurements and the trigonometric functions is foundational in solving different problems that are about these vectors and their orientations. For instance, this angle of sine, cosine, you know, what is this idea of tangent, they are very important. Just to give you an idea, the um uh tangent is specifically used as part of the activation functions, we call it tanh activation function, and knowing this tanh will help you to understand the activation functions that are used as part of deep learning, which are more advanced machine learning type of models, and they are fundamentals in all these different new and cutting-edge techniques, like large language models, Transformers, encoder and decoder based algorithms, etc. They're also important in this idea of computing dot products, so very important and must-know when it comes to linear algebra. So this is just a simple example when it comes to this right angle triangle and Pythagorean theorem and how it is applied. I will skip this for now.

It's also important to understand this idea of orthogonality. So the two vectors, let's say A and B, they are orthogonal to each other if their dot product is zero. So later on, as part of the vectors when we will talk about dot product, we will see what we mean when we say that the dot product is equal to zero. And here you can even see that if the A norm, if the A vector, so you see here and B vector, if those vectors, if we multiply them to each other, their dot product is equal to zero, it means they are orthogonal. So this angle that they form is equal to 90°. Orthogonality implies that the vectors form a right angle with each other. They are in, you know, we, we are dealing with that in R<sup>2</sup>, in R<sup>3</sup>, so they are super important when it comes also to visualizing them correctly. This concept is visually represented, all this, you know, vector A and then vector B, and they are perpendicular in the 2D uh coordinate system. All right. So when it comes to the applications of orthogonality, orthogonality plays a crucial role in various aspects of linear algebra. It's fundamental.

In defining vector spaces, subspaces, in solving a system of linear equations, later on, when we pass the vector ideas and we go on to the matrices, solving linear systems, so equations with many unknowns, and then we use this idea of reductions or Gaussian reductions, we will see how this idea of orthogonality can be important and how also it relates back to the norm of two vectors. So it's fundamental in defining all these different identities and solving systems of linear equations. Also, orthogonal vectors are used in finding the shortest distance from a point to the plane, something that is important when it comes to optimizations.

Here you can see an example: the vector A, which is equal to (2, 3), and then vector B, which is equal to (-3, 2). You can see that when we multiply 2 by -3, so we obtain basically the dot product—by the way, this is something that we are going to cover also as part of this course—but for now, you can see that if we take this number, we multiply with this, so 2 * -3; we take this number, multiply with this, so we take three and multiply with two. You can see that this is equal to -6; this is equal to 6. So -6 + 6 is equal to zero. So you can see that the dot product of these two vectors is simply equal to zero, and this is what we are referring to as orthogonality. This means that these two vectors form a right angle, where we see here this angle is equal to 90°.

Why these prerequisites matter and why I mentioned those: understanding this concept is very crucial. They underpin this geometric interpretation of linear algebra; they will help you to better understand these concepts and not just to memorize them, but really understand. Later on, when you go into your machine learning and AI journey and in your data science journey, seeing these concepts will help you to better understand those different algorithms, these optimization techniques, what we mean when we say we want our optimization algorithm to move towards a local minimum, a global minimum. But this idea of movement, this idea of vectors, later on, you will also understand these different concepts in deep learning, how these models work, how the neural networks work. Those are essential concepts that you need for solving different systems of linear equations, a core part of this course. They also help you in visualizing vector spaces, which are critical to understand this concept of linear algebra, the applications of linear algebra when it comes to real-world applications. So those are things that you can definitely—must—by following some of our other courses. But for this course, I assume that you are already familiar with these concepts, right?

So now we are ready to actually begin, and with these prerequisites in mind, you are prepared to start your linear algebra journey. We are going to learn everything in the most efficient way, in such a way that you will learn the theory; you are going to see many examples; we are going to learn everything in detail; but at the same time, you're going to learn the must-know concepts, and I'm not going to overwhelm you with the most difficult concepts that you will not be seeing in your career. I'm going to give you the bare minimum when it comes to really knowing and the must-know for linear algebra, such that you will be ready to apply linear algebra in your professional journey, whether you want to get into machine learning, deep learning, artificial intelligence, data science. Knowing these different concepts in linear algebra, you will be a pro in your field. I'm going to give you everything that you need: the theory, examples, implementations, everything in detail, but at the same time, you will be doing that in the most efficient and time-saving way. So without further ado, let's get started.

Let's now quickly define this idea of norm. So the norm of a vector, denoted by ||v||, which you can see kind of like similar to the absolute value from pre-algebra, you can see here that we have these double straight lines like from absolute value, then we have the name of the vector or the variable name that we are assigning to our vector, and then you might notice here on the top of this, this arrow; this basically says that we are dealing not with just a variable, but really we are dealing with a vector. This is really important because you can see that there makes a huge difference if we have, for instance, just V or V1, I have to say, or just V; those are really important things that you need to keep in mind when it comes to linear algebra and trying to differentiate vectors from a point. You will notice that when it comes to norm, we can represent it either by this notation or this; usually it's a common notation in machine learning or in data science with these two bars. And when we do this, we automatically also know L2 norm, and this is something very common and usually used as part of regression, which is an application of linear algebra, and it's used in regularization. So we are regularizing our machine learning algorithms. So when you get into machine learning, you will see time and time again this notation. So next time when you see this, then you know automatically that you are dealing with L2 norm. L2 norm, which is also used a lot in machine learning, it is referring to the usage of L2 norm in regression, and regression or L2 regularization is a very popular regularization technique as part of machine learning. So right now, even you can see this intersection of linear algebra or this idea of norms in machine learning.

The norm of this vector v is equal to √(V₁² + V₂² + ... + Vₙ²). So here, basically, it means take the square root of V₁², then V₂², plus V₃², blah blah blah, plus Vₙ². So basically, take all the units that form this vector, and then so are on this vector and use them, square them, and then add them, and then take the square root of that; that's the distance, or I have to say, the norm of this vector. We saw already the norm here is just an example of what norm is on a specific two-dimensional vector. When we have, for instance, that the vector is equal to (3, 4), which means for the first dimension, let's say on the x-axis, we have three, and then on the y-axis is equal to four, then the norm or the Euclidean distance, so this is equal to—we take the x value, so three, and then we square it—so V, you can see here this is the case when n is equal to 2; this is simply equal to √(V₁² + V₂²). And as V₁ is equal to 3, so this is our—maybe I can make this just V₁, and this is my V₂—then the norm or the Euclidean distance for this vector, so this thing is equal to V₁² + V₂², which is equal to 3² + 4², and this value is √25 and it's equal to 5.

Let's now see the difference between Euclidean distance and the norm. So you could see here the norm; here we have just one vector like here, and this norm it has just two corresponding values in two-dimensional space. You see here we have just three and then four, so this is V₁ and V₂. When it comes to the Euclidean distance, this is kind of the generalization of this idea of norm. So the Euclidean distance between two points A and B in Rⁿ, so in the n-dimensional space, is the norm of the vector connecting A to B. So we see that the norm and the Euclidean distance are highly related to each other. Only we are talking about the norm when it comes to one vector, but when we have this vector A and the vector B, this is simply the Euclidean distance. So for the Euclidean distance, we know already this idea of distance, how we can measure it, and you can see that this comes very similar to what we see here, notation. And here we are saying, well, we have this vector, and then it has these two coordinates in n is equal to 2, in two-dimensional space. When it comes to the Euclidean distance, Euclidean distance helps you understand what is this distance between two points in an n-dimensional space. So the Euclidean distance between two points, let's say A and B, in n-dimensional space is the norm of the vector connecting A to B. So for instance, if we have a point A and we have a point B, we are connecting this, and this is the vector connecting these two points, then the Euclidean distance is simply the norm of this vector. So this is the Euclidean distance. So we can see that the norm and the distance, they are highly related to each other. In the Euclidean distance, we're using this idea of norm, and specifically the norm 2, as I mentioned before.

So here you can see that the definition of Euclidean distance: the distance between A and B, the two points, is equal to √((A₁ - B₁)² + (A₂ - B₂)² + (A₃ - B₃)² + ... + (Aₙ - Bₙ)²). So here what we mean basically is that if we have two points, here is A and here is B, and this is the vector, and we know all these different points, so A₁, B₁, A₂, B₂, A₃, B₃, blah blah blah, and then here Aₙ, Bₙ, we know all these points lie here in this distance, then we are taking them and using them to calculate the Euclidean distance. So here, for instance, if we have a point A and B, so in this example, let's do a quick, one specific example when we have a point A which has coordinates (1, 2), so this is basically A₁, A₂, and then point B with two points in it, like B₁, B₂, you can notice that the d(A, B), so the distance or the Euclidean distance of these two points, which is equal to the norm of this vector—or here this is A, and this is B, and this is this vector—this is equal to √((4 - 1)² + (5 - 2)²), takes the square root of that, and says this is equal to 5.

Now you might be wondering, but hey, why do we do then, instead of (1 - B₁)² we do (B₁ - A₁)²? And the answer to this question lies in the properties that we learn as part of pre-algebra, because it doesn't matter when we take (A₁ - B₁)² or (B₁ - A₁)² because this squared ensures that it doesn't matter which one we take first and subtract the other. Now the proof of that is outside of the scope of this course; this is part of pre-algebra, but I just wanted to put this out there to ensure that you are seeing what we are seeing here, because here it says (A₁ - B₁), but in this example, we are taking instead of that B₁, and we are subtracting A₁. This is a common thing that we do in pre-algebra and just in general in different Euclidean distance or distance-related cases. So I just wanted to put this here to ensure that later on this is something that can be clear from the first view, why this is important.

This idea of norms and Euclidean distance, beside being used in machine learning, and why is it used? So norms, they provide a way to measure the size or the length of a vector in vector spaces, which means that when we want to measure a distance, a similarity, a relationship between, for instance, vectors, then it becomes much easier to use this idea. An Euclidean distance is not only used in regularization techniques like L2 regularization or regression, but it's also used in other machine learning or deep learning algorithms as a way to measure the distance or the relationship or the similarity between two different entities. Those can be variables; those can be two people that we want to compare in our algorithm or two entities. For instance, the norms or the Euclidean distance, they are also used as part of the k-means algorithm, something that you might have heard. And if you follow later on the machine learning and the clustering section of machine learning, you will see that Euclidean distance is used as part of k-means algorithm that aims to cluster observations into different groups. So this is also yet another highly applicable topic that you must know in order to understand different linear algebra topics, but also machine learning topics.

Welcome to the course on the fundamentals of linear algebra. My name is D. Vasan, and today we are going to start with some basic concepts that are important for understanding linear algebra. Linear algebra is one of the most applicable areas of mathematics. It is used by pure mathematicians that you will see in universities doing research, publishing research papers, but also by the mathematically trained scientists of all disciplines. This is really one of those areas in mathematics that you will see time and time again appearing in your professional life if you want to become a job-ready data scientist, or you want to do some hands-on machine learning, deep learning, and AI stuff. But also, linear algebra is used in cryptology; it is used in cybersecurity and in many other areas of computer science and artificial intelligence. So if you want to become this well-rounded professional, you want to go beyond using libraries, and you want to truly understand the mathematics and the technical side of these different machine learning algorithms, from very basic ones like linear regression to most complex ones coming from deep learning, like architectures in neural networks, how the optimization algorithms work, how the gradient descent works, and all these other different methods and models, then you are in the right place because you must know linear algebra such that you will understand these different concepts from very basic ones to most advanced ones in the data science, machine learning, deep learning, artificial intelligence, data analytics, but also in many other applied science disciplines.

Before starting this comprehensive course that will give you everything that you need to know about linear algebra, first I'm going to tell you what we assume that you already know, because linear algebra, it comes from about the third year of bachelor's of different highly technical studies, and here we are assuming that you already know certain concepts. So to ensure that this course focuses really on the topic of linear algebra and that you understand all these concepts really well, for that we need to be able to know different topics. So before we dive into these concepts, let's familiarize ourselves with the basic prerequisites and notations used throughout this course, and you will really need to know this in order to understand these concepts really well, such that instead of memorizing, you will actually just hear me once or maybe twice, and then every time you hear later on or you see it in the papers or in some algorithms, you will recognize, ah, this is something that we already learned.

So some key prerequisites overview is here. First of all, to fully grasp the upcoming material, you should be familiar with some basic concepts like real numbers, vector spaces. So you don't need to know this idea of vectors, though you already most likely are familiar with this, given that you know how to plot different lines, you know the idea of x's and y's and how to plot these different graphs. But here we are going to touch base on this. Every time when we come close to these concepts, I will refresh your memory, and we will go through these numbers, the idea of norms and distance measures. Because when it comes to the vectors, when it comes to the magnitude and all these different topics that we are going to discuss as part of linear algebra, knowing what a norm is and what is the definition of distance, what is the length between two points when we plot it in the two-dimensional space or three-dimensional space, those are all very basic concepts that usually you see as part of a basic pre-algebra or is common algebra questions and lessons.

In order to truly understand what linear algebra is about, to understand the direction of vectors, the angle, and then the dimensionality reduction, how linear algebra is applied, for instance, in different algorithms in machine learning, deep learning, data science, statistics, you really need to understand this Cartesian coordinate system. So this is not only important for linear algebra, but I assume you already know it, given that you have passed those other courses like calculus, or usually they are covered as part of pre-algebra or algebra. So the Cartesian coordinate system, I mean here, understanding what is, for instance, the common description of them; for instance, when you, when we write like X and then Y on the vertical axis, and then we can, we have here zero, and then we can always plot these different plots. You know, we have a clear understanding what this Y = X line is; we understand how, by knowing certain points, we can plot different plots; for instance, that this is the Y = X line, that here it means that if we have here one, then this is just one, two, this is two. So we understand when we have the function of the line and we have a certain value, where is our y-coordinate or x-coordinate, then the corresponding coordinate can be found. Then you also need to know some basic things that I just didn't mention right now. So for instance, that the numbers here can be like 1, 2, 3, up to infinity. So you understand this concept of infinity, and then here the same story; then here we have -1, you know, -2, and then this is then used later on, and we will be touching base on this when we will be describing our vectors and how we can visualize our vectors, either two-dimensional space like we have here, because this is two-dimensional, so we have X and Y, but we can also, of course, visualize it in three-dimensional, etc. So this idea of a basic coordinate system is really important, usually covered as part of algebra, if not pre-algebra.

Then we have basic trigonometry, which means that you need to have a clear understanding what sine is, what cosine is, what tangent is, and their reciprocals. And here I mean that you know, for instance, what is the cosine function, what is the sine function; you know that you have an understanding, for instance, that what is this line, you know, whether it's a sine line or cosine line; you have also an understanding what this π is. One thing that I didn't mention, but it just goes around all these topics, some basic things that you understand what is X, what is Y, why we use them, and this idea of variables, and also you need to understand this idea of a square or you know a 90° angle, and then Pythagoras' theorem. Here we have the same, so what is this relationship between different sides of the triangle that is a right triangle and that has one of the angles as 90°? And this idea of, you know, the sides, how this relates to the sine, cosine, tangent, cotangent, and also how the Pythagorean theorem applies when we have a triangle, but it is no longer with an angle that is 90°; what is the sum of all the angles of a triangle? So those are basic stuff that are commonly covered as part of trigonometric lessons or part of general geometry.

Then another prerequisite is this understanding of identities and equations in trigonometric lessons, something part of which I already covered, and this goes around having a basic understanding of algebra and geometry. Those are super important to understand more advanced techniques from linear algebra. Then we have finally this idea of orthogonality, perpendicularity in vectors. For instance, if we have two lines like this, then we are talking about perpendicular vectors. When you have two lines that are actually parallel, so they don't have any intersection, and you won't find any point that is common for the two.

Let's get started with our first module, which is Foundations of Vectors. In this module, we are going to talk about fundamentals of

Linear algebra: vectors. We are going to make a differentiation between scalars and vectors. We are going to define them. So first, we will learn the theory, then we will implement them into practice by plotting them, by looking into different examples. Then we will look into this representation of vectors by looking into the magnitude and the direction of it, and the representation of them just in general. We are going to plot them in our coordinate system. Then we are going to see the common notational vectors and indexing of them.

Vectors are super important when it comes to linear algebra and application of it. And uh, they matter not only in mathematics but beyond. So uh, vectors help us in many ways, from figuring out how objects move to solving math problems in science and just in general in technology, including in data science, machine learning, artificial intelligence, etc. They are a super useful tool. So uh, let's start our journey with looking into scalars.

Scalars: they are just plain numbers. And by definition, a scalar is a single numeric volume, often representing magnitude or quantity. For example, uh, scalars can be describing um the temperature outside, for instance, the temperature of um a 22° uh can be represented by a scalar, or a height of a person can be represented; it's a scalar. So let's assume we have a scalar that we will define by a letter s; it's just a variable. This scalar is then equal to 22, for instance, and we are measuring it in degrees. So it means that uh, if this s measures a room temperature, then the scalar s, which is equal to 22°, which represents the room temperature, it can be for instance 18° or 9° if it's very cold. Uh, it just measures a single volume; it represents just a single number, or it can be for instance 17,100, 2.22. So all these, they are just scalars; they represent a single numeric volume. They often represent a magnitude or a quantity. We will see that scalars, they are a value that represents the magnitude of a vector. So uh, now when we are clear on this very basic concept of scalars, let's actually move to this idea of vectors.

By definition, a vector is an ordered array of numbers which can represent both magnitude and direction in space. So uh, vectors, they are a bit more; they represent a bit more than scalars. There are numbers that also show direction, like a car speeding down the highway or a bow uh being thrown, for instance. Uh, when it comes to our previous example, we were using this uh uh room temperature as a way to uh think about the scalar. A scalar, for instance, a scalar that we just saw was this room temperature, room temperature, which was 22°. When it comes to the vector, a vector is different. For a vector, for instance, we can have an example when a bird, for instance, a bird, it flies flies at 10 kilometers per hour, and I also add here another information which will make this as a vector, which is that it flies South. So here, as you can see what I'm doing is that I'm not just—oh, let me actually remove this part to make it easier to understand. Okay, so uh, in this example, let me write it down that the example: bird flies at 10 kilometers per hour. So you can see that I'm not just adding the scalar, which is in this case the magnitude—we will see very soon the formal definition of it—so I'm writing down the speed; I'm defining the speed, but also the direction. So I'm saying I know that the bird is flying South; that's the direction, and I know also the speed of it, which is the magnitude, so 10 kilometers per hour. So here in the vector I have much more information than in the scalar, because in the scalar I just got temperature, room temperature, single value, but in case of a vector I not only have um magnitude or speed, like 10 kilometers per hour, but I have extra information, which is the direction of it, for instance, flying to the South.

Let's now look into some real examples and plotting them to make more sense out of this idea of vectors and what is this magnitude, what is the direction. So let's assume we have a 2D plane. So we have x-axis, we have y-axis here, like usual, we have our z0 center, and we want to plot a simple vector. So uh, usually the way we represent a vector in tutorials or just writing down is by writing the name of the vector; this can be just a a random name. Let's assume that it's a v, letter v, and then on the top we are always adding this arrow. So this arrow, it says and it tells the person who is reading that we are dealing with the vector; arrow on the top is that reference. So let's assume this uh vector v, it starts from the center of our coordinate system and it goes to this point. So let's say in here, this is our vector v. So let's assume that this point in here is equal to 4, which means that the x-coordinate is 4 and the y-coordinate is 0, as the um uh arrow, it just as the point in here, it has a a y value of 0. So you can see that it goes straight from 0 to this one, to this point. Okay. So what tells this vector uh to us is that we have a value that describes the length of the vector, so it goes from 0 to 4, which means that the length is equal to unit 4, so it's equal to 4. Um, and we have just learned and we were just talking about that the magnitude is the length in this case. So the length describes the magnitude in this case. So this means that the magnitude of this vector is equal to 4. And then um what else we can see here? We can see the direction of the vector, which means that the direction is also something that we can see here; this is the direction of the vector. So this going straight from this point to this point in a horizontal way. So independent whether I plot this vector from 0 to 4 in here or in here, here or in here or in here or in here; in all cases, as long as the length is this, I'm dealing with the same vector, because I am basically in this entire R2 space; I have exactly the same vector. All I care is about the magnitude and the direction. Where will this vector start and where will it end? I am not interested; I'm interested that the uh that the magnitude, in this case the length, is equal to the direction of the vector.

Let's now look into another example where we go a bit more difficult on our coordinates and on our vector. We already saw that we had this vector where we went—let me change the color—so this was our vector v, and it went from zero till 4. So this point, to be more specific, is—so this vector, it goes—the vector v, it goes from 0, 0 to 4, 0. So the coordinate x was 4 and the y was 0. Now let's plot another one um where the direction is no longer horizontal. For this vector, let's call it vector w, and for this vector w we will again start with z0, so we will start again in here, but this time we will go bit like this. So let's say we go all the way to this point. So this point has a value for an x-axis of 3 and for y-axis it has a value of 4, which means it goes from this point to this point, and this is the direction of our vector v. So it goes to 3, 4, because this point is 3, 0 and this point is 0, 4. So x-axis is 0, x-coordinate and y-coordinate is 4. So now you can see that the direction of this vector is like this, while the direction of the vector v was like this. And like in case of vector v, I again no longer care about where exactly my vector w starts and ends, but all I care is about its magnitude, so the length and the direction. So for instance, I can have the same vector in here, the same vector in here, as long as the length, the magnitude is the same and the direction, I am dealing with the same vector; that's all I care. So the magnitude and the direction is all that you care about. All right. So now about the length, um that's uh something that you can see very easily from this specific example, because by using the Pythagoras Theorem or Pythagorean theorem, we can see very quickly that as the length of this side of our uh right angle 30°—so right triangle—we can see that this side is 3, this side is 4, which means that this side is 5, because 4^2 + 3^2, then we take the square root of that; square root of 25 and it's equal to 5. So the length or the magnitude of this vector v is simply equal to 5. All right, this was about this uh specific vectors. Let's now look into the uh common representation of the vectors.

So we always use the magnitude as well as the direction, you know, to represent the vectors, and they commonly are represented by two different uh ways. Let's now look into the first way that the vectors can be represented, and then we will move on to the next one. So when it comes to the vector v, so we saw that vector v was moving from 0 till uh to the point of 4, 0. So we can represent the vector v by (4, 0). When it comes to the vector w, we can represent that uh vector—so vector w, we can again do the parenthesis and we can say that it's equal to (3, 4). So by using the coordinates from the coordinate system, we can then represent our uh vectors. So this is just one way of representing a vector. Another way of representing these vectors is by using these square braces. Given that we are in a two-dimensional space, first we will mention here the 4, then we will mention the 0 in here, two. So we can say [3, 4]; this is yet another way of represented the vectors in a two-dimensional space.

So if we were to have a three-dimensional space—so let me actually show it on a new page—so if we were to—if we were—to have vectors in three-dimensional space, so we are dealing with R3, so we have points that can be described by x, y, and z, so coordinate space like this, so x and and the y and then the z, then every point—so let's say we have this vector—then we had to represent it by a value, let's say x, x1, y1 and z1, or um better let me actually use different letters, a, b, and c, and this would be my vector v, and I could also represent this vector v as the same—so vector v can be represented as (a, b, c). So one thing that you can notice is that unlike the R2, now I have three different entries, what we are also referring as rows, and we just got one column. So um we can uh often represent and usually that's a common way of representing vectors by using this um columns. Columns help us to represent our vectors, and you can see very clearly then when it comes to the two-dimensional space, so when we have R2, so then our vectors have just two rows, so [3, 4], [4, 0], like in here. When it comes to three-dimensional space, we have three entries and so on. So the same holds of course also for for instance R5, then for R5 um our vectors—so coordinate space can be for instance x, y, z, and then γ, and then let's say δ, and then the coordinates uh of a vector in that space can be v and then arrow is equals to, and then we would have uh let's say (a, b, c, d, e). You get the idea. So depending on the space, the coordinate space and the dimension of that space, then the corresponding vectors can be represented accordingly.

So the vectors are quantities that have both magnitude and direction, as we just saw, distinguishing them from scalars which only have magnitude. So we saw that the scalars got only magnitude, while in case of vectors we saw both for the vector v and for the vector w; we didn't we didn't only have the magnitude, so the length of the vector, but also the corresponding direction. So uh, when it comes to the um vectors, so this is exactly what we just saw in our example: a vector in a two-dimensional space, so in R2, uh can be represented by using this square braces and the corresponding entries for x and y, where x is basically the x-coordinate in our coordinate system, so in our x and y system. Whenever you have this x and y coordinate, then uh this x coordinate will then describe your magnitude and the y coordinate will then describe your second entry that you need to put when representing your vectors. So here the x and y indicate the movement in the horizontal and in the vertical dimensions respectively. So for x's it's always the x-coordinate, so how far you move towards the horizontal direction, in here, in here, or independent in here, so always take the x-coordinate; that is the value that you need to put first, and then the y needs to be put it in here.

Indexing in vectors. When it comes to the um indexing, the standard mathematical notation uh indices in the n vectors goes from i = 1 to i = n. So the um notation here can be a bit ambiguous. So ai uh could mean the i-th element of ai, the a vector or the i-th vector in a collection. So let's start with a simple one and then move move on to this next part. So what this means and what this means we will look into now. So uh usually uh when we have a um n-dimensional space, we are having a hard time visualizing it; therefore, we use this two-dimensional space or maximum three-dimensional space in order to get an understanding of what these vectors are. So we just saw examples of them uh when uh creating our vectors in um v and v uh and w in uh R2 and also in R3, but we can have similar vectors also in R4, in R5 or all the way down to R<sup>n</sup>, where n can be 100, 200, 500, any number as large as you want. The thing is is that visualizing R<sup>4</sup>, R<sup>5</sup>, R<sup>n</sup> is very hard, but we can still benefit from this great properties of the vectors, matrices, and in general linear algebra in order to describe different things that have more than three dimensions. Therefore, we have this a bit more ambiguous notation where we use R<sup>n</sup>, and this n can be any real number and it can be all the way to infinity, so a very large number. And uh, let's say we have a vector in this R<sup>n</sup>, then this vector is usually described by using similar square uh brackets like before, only with uh more entries. So like before we got just one column, so that's something that we didn't uh change, but here we have instead of just two entries or three entries like in the two-dimensional or three-dimensional spaces, now we have a<sub>1</sub>, a<sub>2</sub>, a<sub>3</sub> all the way down to a<sub>n-1</sub> and a<sub>n</sub>. So we got in total n elements in our column, and this describes our uh single vector. So this vector in an n-dimensional space, this we can call also a—.

So one thing that we just saw is that it was saying in our definition and notation that uh we might also be dealing with the i-th vector in a collection, which means that sometimes you will see this, while here the a<sub>1</sub>, a<sub>2</sub>, they are vectors themselves. So here we saw that these are just entries, so a<sub>1</sub> is a number, a<sub>2</sub> is a number, a<sub>3</sub> is a number, a<sub>n</sub> is just a number, but it's also possible uh when you have a much more difficult and complicated case that you got an A—let's write it down with a capital letter A, which is equal to A<sub>1</sub>—or let's actually remove this—so we got let's say A<sub>1</sub>, A<sub>2</sub>, A<sub>3</sub> all the way down to A<sub>n-1</sub> and A<sub>n</sub>, where you can already see what is going on. So instead of having just a number as an entries, instead we have vectors in here. So our first element is actually a vector, our second element is actually a vector, so A<sub>2</sub>→, A<sub>3</sub>→, all the way down to A<sub>n</sub>→. So while here this can be for instance some numbers, let's say 1, 1, 1 all the way down to 1, 1, here we have a vector, vector, another vector, and all the way down here yet another vector, where for instance—let me remove this part—where for instance A<sub>1</sub>→ is actually equal to (a<sub>11</sub>, a<sub>12</sub>, a<sub>13</sub> all the way down to a<sub>1n</sub>). One thing that you will notice here is that unlike in here, here I got double indices, so I got here a<sub>11</sub> and then a<sub>12</sub> and then a<sub>13</sub> all the way to a<sub>1n</sub>. So the first index it doesn't change as I have here a<sub>1</sub>, so I'm writing down the index corresponding to this vector, but the second index it changes per entry indicating which element specifically in the vector I'm talking about. So from the first index you can identify the vector that I'm referring to, which is A<sub>1</sub>, and from the second index you can see the corresponding um entry or the value that that um element is positioned in this vector. So you can see that this values, for instance, in the um vector 1, so A<sub>1</sub> to be more specific, but then it is in the first position, this is in the second position, in the third position, all the way down to the n-th position.

So this is something that is really important to understand well, because this notation is going to appear time and time again across various applications of matrices and vectors, so is really important to understand well. Therefore, I want to go one more time through this to make sure that we are clear on what this indexes represent. So whenever we have an index uh an a vector that we want to uh represent and it's um it has just um it is just a vector, which means that it's not a nested vector, vector in a vector, then um we can define it by, let's say a, and then on top an arrow, and it's equal to, and here we can have a<sub>1</sub>, a<sub>2</sub> all the way down to a<sub>n</sub>. So you can see what we are also referring as dimension of this vector is equal to n x 1. So I got n entries and just one column, so n x 1, which means that this already gives me an indication that most likely this a<sub>1</sub> is a number, this a<sub>2</sub> is a number, this a<sub>n</sub> is a number. So let's say this equal to 1, 2, uh 3, blah blah blah, and then here I have let's say 100. But if I'm dealing with the nested vector—later we will see that this can be represented by a matrix—then um I can also define this by capital letter A, which is a common way to refer to either matrices or nested vectors, and then this is equal to A<sub>1</sub>→, A<sub>2</sub>→, A<sub>3</sub>→. This already sends a message to the reader that we are dealing with no longer u constants within a vector, but rather vectors in a vector. And uh what can we see here is that the dimension of this nested vector, or which we can also refer to as a matrix here, the number of rows, so the number of entries, this elements, we can see it's equal to n, but then this time the number of values that form these vectors is no longer one, because we are not dealing with just a constant; this is not some constant, but rather this is yet another vector. So let's assume this vector has a length of m, so let's say this has a length of m, then the dimension of this matrix a is equal to m. So something that we will see also when talking about matrices—so let me actually clarify this bit more for better understanding—let's say we look into one of those um one uh one other example of an entry, so let's say we look into this specific vector which is in the uh the third uh vector within this vector capital A, so this A<sub>3</sub> vector. So one thing to see here already is that I assumed that these vectors they got m elements, and keep in mind that all these vectors they should be of the same size, so it means that I already know that this specific vector A<sub>3</sub> has m elements, so m elements. So I'm representing this uh A<sub>3</sub> vector—from here I'm taking this out from this entire uh nested a vector and I just want to represent this—and now unlike this elements that got an arrow on the top—

This time, I will have uh, constants forming the A3 Vector. So I no longer have vectors, but I have elements in it. So in here, I will have a, a, let me actually write down all the A's, but to refer and to make sure that I recognize that I'm dealing with the third A vector, so A Tre Arrow here, I will put three Tre all the way here Tre, so they all come from the same third A3 Vector, but then their positions is different because this is, let's say uh, one, two, and then all the way down to Ed position. So this indices help us to keep track what are the um, position that these values are taking part in the vector A3 arrow.

This might seem bit complicated at the moment, but once we move on onto bit more complex material like uh, matrices, it will make much more sense. This is bit of an extra; I just wanted to Showcase this, but this is what uh, is at its core and what you need to uh, understand at the moment to understand this concept of vectors. So you need to know that vectors can be represented by this arrow on the top, so let's say Vector a, and it has let's say n elements, then you can write the square brackets, and then you will need to mention A1, A2, all the way to a n, which means that you have n different entries describing your vector. So you have A1, which is the first element in your vector, A2 the second element, all the way to a n, which is the end element. Where here you can see, for instance, so if I had here A3 that uh A1 is simply equal to one, A2 is equal to 2, A3 is equal to 3, all the way to a n is equal to 100. So this numbers I'm basically taking and I'm representing them, I'm putting them in here within Square braces in order to get a representation of my Vector. So my Vector a has all these different entries and different entries, and it starts with one and it ends with 100. This is a vector, and then when it comes to the vectors within vectors, here we need to be a bit more careful. CU here we not just have uh, constant values forming a vector, but we have vectors that form yet not vectors. So our Vector a, our nested Vector a, which we uh later will refer as Matrix a, has actually entries that also are vectors. So we have a 1 Vector, A2 Vector, A3 Vector; they are not just constants, but only own they are vectors.

So here, for instance, we have defined also an example of it; we have said let's look into this third specific Vector that is part of a, which is A3 uh, vector, and uh, that one has M different elements. Here we have then the index referring to the which Vector from the nested Vector a it is, which is the third one because we have taken it from here, but then on its own this Vector has different members and different members to be more specific; therefore, we have also an index to keep track of the position of this value one to up to M, and this can be yet another uh, this time it can contain some elements, an example of which is, for instance, Z 1, 2, all the way to let's say 500, and this can be different numbers; it doesn't need to be ordered, it doesn't need to have a specific pattern; they can be just random numbers describing this A3 Vector. So hopefully this makes sense; if it doesn't, don't worry, because we are going to see this time and time again. I just wanted to give you a brief of an intro such that you can uh, remember this when we come uh, back to bit more uh, complex topics like uh, indexing in matrices.

So now let's talk about special vectors and operation. Here we are going to talk about zero vectors, unit vectors, the concept of sparcity in vectors, as well as vectors in higher Dimensions like we just saw about this n dimensional space. We will also talk about different operations we can apply when it comes to vectors like uh, addition, subtraction, and then later on in the next module we will also talk about multiplication. We will also be looking into the properties of vector addition after we have looked into some detailed examples when it comes to operations on vectors. All right, so let's start with the zero vectors and unit vectors. When it comes to zero vectors, you can see here already that um, the zero and arrow on the top it basically refers to the vector like we saw before, only with the difference that all its members are zero. So you can see here that we have zero and then an arrow, and then underneath here we have some number tree, and then this is described by this common representation with the square braces and then three different members z0, 0, so all zero, and then it says in R Tre. Okay, so why are we doing this? Well, uh, when it comes to uh, different linear Lal operation, sometimes we just need to add zero vectors, or we just want to create zero vectors; it's just easier to work with. You want to uh, just create an empty uh, Vector; we want, we know the length, but we want to keep it empty such Laton we can add something on the top, or knowing that when we add a zero on a number the number stays the same, we can make use of this property to uh, do different um, uh, tricks when it comes to programming in Python, in SCAR, or in C++ Etc. So therefore, this idea of zero vectors can become very handy.

Now, one thing that you need to notice here is that we are not just writing down this zero to emphasize we are dealing with the vector, but like uh, before we have this error on the top emphasizing that we are dealing with a vector, then what we are doing is that we are also adding the dimension of this Vector. So in what dimension, in what space are we um, uh, creating this zero Vector that this Vector is located? Is it in r R2 in RN in R3? In this specific case, you can see that in this example the uh index that we got here is three, which basically indicates we are dealing with a zero Vector in threedimensional space, so in the R3 uh, in general we would just note this by n, keeping the uh notation general, which means that we are dealing with 0, 0, all the way down to zero, so it has n one dimension in r n. All right, so this is about zero vectors; it is just a way to uh, make our programming life easier, also to use it in different uh, algorithms when it comes to bit more advanced algebra.

The next type of special vectors that we will look into is this unit vectors. So vectors with a single element equal to one and all the others zero, denoted as EI for the E unit Vector in N dimensions are referred by unit vectors. So uh, what we mean here when it comes to the unit vectors, um, if we have for instance E1, it means that we have a vector where the e, in this case the first element is equal to one. So you can see that E1 is equal to 1, 0, 0. So in the first element we got one, and the remaining is zero, and this is really important that we are dealing with vectors that contain only elements of zeros and ones, and the only member that is equal to the only element in that Vector that is equal to one is the E element in the entire Vector; all the remaining ones are zero, and you can see here that the dimension is no longer specified, but just the um index of the entry where the um uh, the uh one is located. So let's look at another example in here, for instance, when it comes to the um uh, unit Vector, yet another unit Vector is E2, which basically means that in the second element, so in the second place uh, the uh Vector contains one, and all the other members are zero. So here you can see Zero here; it can see Zero only in the second element we have one, and then in the E3 what we have here is that the third element is one, and all the other ones are zero. So let's actually look into uh, one um, bigger Vector uh, in higher Dimension to make it even more sense. So first I will Define and assume that we are dealing with a vector in RN, so in an N dimensional space, this gives me an idea that we are dealing with um, so we are not dealing with nested Vector; we are dealing with a simple n dimensional Vector, so it has n rows and one column. So using the square braces I'm going to represent my Vector, so I have all these different members n members C, so e, let's say it is E5, so what does this mean? It means that I is equal to 5, and this I element, so the fifth element is equal to one, and all the other entries, the elements in this Vector are zeros. So let's look into this is z0, 0, z; I'm approaching the fifth element in my Vector, so it's this one; this is one, and the remaining all zeros. So this is a unit Vector in an N dimensional space, and I'm defining it by E5 because my fifth element is equal to one.

Now, those are very handy when it comes to some other uh, techniques in linear algebra and just in general; think about techniques like um, uh, row etum form, solving linear equation, something that we will see as part of the next unit. So many things um, we can do by using unit vectors; unit vectors are super important, so you need to understand this concept uh, very well such that later on you will understand uh, more advanced concepts in linear algebra. Let now look into the topic of sparsity in vectors. So by definition, a sparse Vector is characterized by having many of its entries as zero. So its parity pattern indicates the position of a nonzero entries. So uh, what we are basically saying is that if we are dealing with a vector that contains too many zeros, we are dealing with the sparse Vector. So uh, this sparsity pattern indicates uh, all Al the positions of a nonzero elements. So um, if we have um, unit Vector, it means that we are already dealing with a sparse uh, Vector. This is a concept that is super important when it comes to linear algebra, but also in general data science, machine learning, and AI, because having a spity in your vector it means that you don't have much of an information; usually a value zero it means you don't know much about that specific volum, and if you got just too many of zeros and too few numbers which do um provide information, it means that you are dealing with a vector that doesn't provide you much information, and there's always a problem when it comes to data science, machine learning, and AI. So sparcity is something that you need to be aware of; you need to know how to recognize it, and you also need to know whe there's a problem in your specific case or not.

So let's look into an example. Let's say we are dealing with this Vector X that has five different elements. So X is a vector coming from um, five dimensional space. So we have for instance an element of three, the first entry, then we have z0 in the second and third uh entries, then we have an entry um four, which coincident also contains value four, and then the last element in our five dimensional Vector X is equal to zero. Now what do we see here? We see that the majority of elements of a vector X is equal to zero because we got in total five elements, and then we got three of it actually uh being equal to zero, and only two of them containing information like equal to three and four. So only two elements that are not zero, so non-zero elements, it means that 3 / to 4, which is basically 60%, 60% of all the entries in the vector X are equal to zero. So the 60%, it means that is above half, so above 50%, 60% of all the information in this Vector um, the majority is simply equal to zero. This type of vectors we are uh calling sparse vectors, and sparcity is really important concept uh that we need to keep in mind later on.

So uh, while we can visualize vectors in two and three dimensions in linear algebra like we just saw in case of this n dimensional vectors, visualizing uh, the this type of higher dimensional vectors becomes very difficult. So uh, this mathematical flexibility uh, to work with uh, this type of uh, information, so when we can represent uh information many with many entries, we can represent it by vector s which we can actually not visualize becomes very handy for complex data structures, for different simulations in physics and much more. So uh, we just saw in couple of examples uh, how we can represent vectors in a high dimensional space using this Square braces and this common Vector notation representation. We saw that in an N dimensional space we could uh, very easily represent this uh, very large Matrix or vectors uh, by just um, using this Vector representation. For instance, if we got a vector that had any different entries where n is for instance thousand, so let's say we have thousand, then uh, we can represent uh, this uh, vector or this information by using common Vector notation, so A1, A2, all the way to a th000. So of course we cannot visualize this; it just doesn't make sense; we can visualize two dimensional vectors; we can visualize three dimensional vectors, but we cannot uh, visualize thousand dimensional vectors, so Vector that comes from uh r, but what we can do is still make use of this very useful information in order to uh, do different operations when, and later on we will see that uh, this property and specifically this part of linear algebra it helps us to work with vectors in any number of Dimensions, whether thousands, million, billions; this mathematical flexibility is super important for more complex data structures, uh, for metrix multiplications when, for instance, we are doing different uh, algorithms including how we can represent uh, very large matrices, very large feature spaces; all this different information we can represent just by making use of vectors coming from this specific uh, part of linear algebra.

Let's now finish of this module by looking into some applications of vectors. So one common application of making use of vectors is uh, when we are performing different operations while having words and we want to count those words. So this is a super common application of vectors, and we can account this words, and you can even plot a histogram over how often each of these words appear in a document. So a vector of a length n, for instance, can represent the number of times each of these words in a dictionary of n words appears in a document. So uh, just for the sake of Simplicity, let's assume that we got um, dictionary that contains only three words; of course, in the reality uh, the um, dictionary, what we also refer often as Corpus, it contains much many much more many words, but for the Simplicity we will assume that we just got three different words in our dictionary, so that's a total. Now let's assume that we got a document uh, with these different words, and we want to count how many times each of those words that we got in dictionary actually appear in our document. So uh, if our document is described by this Vector, so it contains an entry of 25, 2, and zero, it means that in our our document we got 25 word one in now from our dictionary, so in the position one, two * word two, and zero * word three. So basically we have a predetermined set of words in our dictionary; in this case three words, word one, word two, and word three, and they have a specific IND specific position in our vector, and when we are putting these values in here, then the machine or the uh computer, the program will understand that if we have 25 in the first position, then the word one in the dictionary appeared 25 times in our document, whereas the second word appeared only two times, and the last word for three didn't appear at all, so zero times in the entire document.

So let's look into a practical example actually to make even more sense. So um, this is by the way a common practice to count different variations of a word; there are common application in engrams, large language models, Transformers; they are just the Cornerstone of many language models when we want to count the words in the document to understand how often the word appears, because this gives us a idea what this document is about; knowing how many times the same word appears in that uh document, it gives us an indication of the topic of uh, the do document. Also, we can make use of a to do sentiment analysis to understand what this document is about, not only in terms of the topic, but also is it a positive, is it a natural or a negative uh, document so to say. So uh, for instance, if we got uh, the following words uh, that correspond to our dictionary, and in our dictionary we got just um, let's say 6 different words, then what we can do is that we can say 3, 2, 1, let's say Zer, 4, two, and the corresponding words are word, row, [Music], number, horse, is, and then document. What this means is that we have a text what we refer as a document that contains three times the word word, that contains two time the word row, contains one time the word number, zero times the word horse, and four times the word eel, and two times the word document. So uh, this is basically a common way representing the uh frequency of the words in the document. Let me actually give you uh another example, and in here I want to emphasize another thing, the concept of stop words. So uh, let's say I make this 10, and then here I say there is a three times the word I, two times the word uh, reading, two * the word library, four times the word book, 0er * the word shower, and 10 times the word uh. So uh, you can see a that in here we are dealing with the document that contains 10 times the word uh, which is what something that we refer as a stop word. So those are things that actually don't give us too much information about what the document is about because uh, it's just used everywhere, but it is appearing too often. So you can see 10 times the most frequently appearing word; this is what we refer as a stop word, and then another thing that we can observe, the second thing we can observe is that we are dealing most like ly with a document that describes library, reading uh, because you see the words like reading, you see the word like book, library, but another word shower that is totally unrelated to reading, book or library is appearing zero times. So even by looking at discounts we can already get an idea what a topic of this document is about. So uh, you can see already know from this very basic example where I made too many assumptions regarding how small the the uh dictionary should be uh, you can even see now how we can use discounts in our dictionary from our text in order to get idea about the topic of the document or topic of the conversation; it can be topic of the uh tweets if you have a tweet data; it can be topic uh regarding book if you have many book um, uh book text; it can be for instance the topic of the review if you got a reviews from uh Amazon, for instance; using this count can help you to get a topic regarding topic from that text; then you can also use it to remove the stop words, because usually the stop words are the most frequently P words; it can also give you an idea about the sentiment; for instance, here we are dealing with natural sentiment; it's not positive; it's not negative; it's just reading a book in library that kind of topic. So all this can be super helpful when it comes to natural language processing; that's a field where this uh text processing, text cing, and then using that for modeling purposes is what uh what plays a central role; it also plays a super important role in the large language models, in the Transformer models, and uh, in simple matters like uh, back of words or uh, in the uh TF IDF; all these they are based on this idea of counting words and how we can use it information, and you can see how vectors come into play in the different applications of linear algebra in data science, natural language processing, in artificial intelligence, in machine learning; so they are super important.

Another application of vectors can be representing customer purchases; for example, an N Vector P, so let's say p can record a customer purchases over time.

With pi being the quantity or dollar value of an item, I now ask: what does this mean? So let's say we have Vector P that represents the customer purchases. We are dealing with a single customer, and we are just saving over time that information: how many times this customer has made purchases over time. The quantity is in, um, dollars; so the dollar value of item I purchased. So we are basically keeping track of, uh, what is the value of the item I that the customer has purchased.

So what we can do is we can assume that in here, actually, it already makes that assumption; it says n Vector, which means that the number of rows or number of, um, items that the customer purchases is n. Now what the, um, the problem says that it represents is that in each entry—and here we have in total n entries—we got a dollar value of item I. Which means that here, if I have P1, P2 all the way to PN, and here somewhere in the middle I got Pi in the ith position, it means Pi represents the value of item I.

For example, if I'm dealing with a customer that buys, um, let's say, uh, courses, and the first item that the customer is buying is a mathematics course—so I'm writing mathematics course—and this is the first course that it buys; e is, by the way, just a, um, way to refer to the ith purchase. Let's say, um, here somewhere in the middle the, um, customer decides to buy a deep learning course, deep learning course, and then it continues buying—the customer continues buying courses—and the last course that a customer buys is, let's say, um, career coaching course.

Now let's say the mathematics course costs, uh, around $1,000. Let's say the, uh, deep learning course costs $33,000, and then let's say the career coaching service, which is usually one of the most applied and personalized ones, can cost all the way to $5,000. Now we see that in the ith position—this is the ith position; let me change the color, by the way—so let's say this is the ith position, this is the first position, and this is the last position. So those are just indices. We can see that in the ith position we got the 3,000, which means that the Pi is equal, equal to $3,000. So this indicates that in the ith purchase the customer purchased a deep learning course, and the value of that item was equal to $3,000. All right. So now we are ready to go on to the next major topic, which is about vector addition and subtraction. So we are going to do some operations and apply these operations to vectors.

So let's first formally define this ideal of vector addition. Uh, two vectors of the same size are added by adding their corresponding elements; the result is a vector of the same size. So, uh, let's unpack this. It says two vectors of the same size are added by their corresponding elements. So here it refers to two different vectors, let's say vector v and Vector W, and it says let's add them—what we refer to as vector addition—and says for that what we need to do is to take all the elements of v and then all the elements of w, and using their corresponding elements—so indices that helps us to understand where those elements are located—we are using in order to add each element in the vector v to the element of the vector W in the same position. Do note that in the second part it says the result is a vector of the same size because we are adding two different vectors of the same size; it's mentioning here it means if we add two different vectors that have the same size, we are going to end up with a vector that has the same size. Now, once I go into the examples, it will make much more sense. Let's quickly also look into this concept of subtraction.

So on its own, uh, subtraction is very similar to this idea of addition. So if we have a subtraction—let's say we have vector v; we subtract Vector W—then we are doing basically, uh, what we just did to the addition, only instead of, uh, doing add, we are doing subtract. So again, we are just—we are just subtracting from vector v Vector W; they have the same size, so we end up having the result, which is a vector of the same size. Only one thing that you can see is that this can be also written as V Vector plus and then minus W. So we basically can represent subtraction, um, on its own as a way of adding, only we take the negative—so the, um, opposite directed Vector. So this will make even much more sense once we go on to the examples. So let's look into our first operation example where we are adding two different vectors. This is a basic example; we got just two-dimensional vectors. We got Vector a that has entries two, three, and Vector B that has entries one, four. And what we are doing is that we are adding Vector a to Vector B. We just learned that we need to have the same size of vectors. So you can see that Vector a has a dimension 2 by 1; vector B has a dimension of 2 by 1; so their sizes are the same—both they got two entries, only two elements—and at the same time we just learned that what we need to do is to take their corresponding elements and add them to each other.

Now what does this mean? It means that we take from a the first element, two, and then we take the first element of the second Vector, which is the B, so we take the two from here and one from here—the first element of a and the first element of B—and then we are adding them to each other: 2 + 1 is equal to three. And then the same holds for the second entry: so three, which is the second element of vector a, and then four, which is the second element of vector B; we are saying 3 + 4 is = to 7. So let me write it down even in a simpler manner such that it will make much more sense. So Vector a has elements two, three; in the first element we got two; in the second element we got three. So a; then we want to add B, which has in the first element an element equal to 1, and the second element is equal to 4. This means that if we want to add these vectors (2, 3) + (1, 4), this is equal to: we need to take two; we need to add one—so this element and this element—and then we need to take three; we need to add to four—so this one and this one—which is equal to: 2 + 1 is equal to 3; 3 + 4 is equal to 7. So we got, uh, Vector (3, 7). Do you note that this Vector, the result Vector, contains again two elements and just one column, so 2 by 1? So you notice that the size is the same of this result Vector.

Now let's actually generalize this concept before moving on to the next example. So if we got, let's say, Vector a that contains n elements: A1, A2 all the way down to An, and it is from n-dimensional space, and we got Vector B that also has n elements—so remember that they both need to have the same size—so B1, B2 all the way to Bn, so they come also—B comes also from n-dimensional space—so then when we add a to B, this is equal to A1, A2 all the way to An plus B1, B2 all the way to Bn. So n by 1, n by 1; the sizes—this is equal to—let me actually use this color to make it even more visible—so I, for the first entry for my result vector, I will get A1 + B1, then A2 + B2, so all the way down onto the nth element, which is An plus—then me use a different color—A1, B1, B2, Bn. So you can notice is now in general terms what we are doing here. So we are taking the A1 coming from the vector a; we are adding in the same, uh, position the value that comes from Vector B, which is B1; we are saying take the A1 + B1; this is the, uh, first element, so the position stays the same, and then in the result Vector. So we take all the corresponding values that are have the same position in the corresponding Vector, first from Vector a and then Vector B; we are adding them, and this forms our new vector, and this new Vector will again have a size n by 1. So you can see that the sizes of the two vectors are the same; both have n elements, and then we are using their corresponding elements to add them to each other element-wise, and then we are getting the result that has the same size, so n by 1. So this is a more general description of how you can add two vectors. Let's now look into this specific example. So we have a vector with the entry (0, 7, 3); so this comes from R3, you can see—so three-dimensional vectors. The second Vector is (1, 2, 0), and then the final result is (1, 9, 3). So how we got this: we took zero; we added 1; 7, we added 2; and then three, we added zero. So you can see all these elements element-wise, and then this is equal to: 0 + 1 is 1; 7 + 2 is 9; and then 3 + 0 is 3; exactly what we got here. So again, the same sizes, and the result is from the same size—so quite straightforward.

Now when it comes to the vector subtraction, what are we doing? That, um—so what are we doing here? So we are doing kind of a very similar thing; we are taking this element one; we are subtracting the other one in this first element; then we are taking the nine in the second position and subtracting this again from the second position of the second vector, and we are putting in here 1; and then 1 - 1 is = to 0; 9 - 1 is = to 8; so we get result Vector (0, 8), like in here, and you can see that the sizes stay the same. So also in this case, let's write more general, um, this idea of subtraction. If we got a vector a from Rn—so n-dimensional space—and it can be represented by A1, A2 all the way down to An, so it has n elements, n by 1, and then we got B also from Rn—so coming from the n-dimensional space, which means that it got n elements—so B1, B2 all the way down to Bn, again with the same size n by 1—then a - B is simply equal to: a—let me actually use the same colors to make it easier to follow—so let me first draw my square braces, and then here I will use blue for a, and then red for the, uh, color for the second Vector, which is B; here I will use black—minus—then, given that the same size should be for the result Vector, I already know that I expect n different elements for this, and then here I'm taking this first element that comes from Vector a; I'm subtracting from this the first element that comes from Vector B, so element-wise subtraction, B1, and I'm already getting the result for the first element in my result Vector. So you can see A1 - B1; I'm taking this element and this element and subtracting them from each other to get A1 - B1, and then the same holds for all the other values, only coming from different elements from Vector a, subtracting from this the corresponding values element-wise from the vector B, so B2, B3 all the way to An. So you can see that in my result Vector, a vector - B vector, in the first element I get A1 - B1, then A2 - B2, then A3 - B3 in the third element, all the way down to the nth element, which is equal to—oh, this should be Bn—so, um, this already should make much more sense. Every time we take the element in the same position from one vector than the other, we subtract from each other in order to get the corresponding element in the final Vector. All right. So let's now, uh, before moving on to the properties, um, I want to show you, um, this only in a coordinate space. So what this means in terms of visualization in a coordinate space. So, uh, let's say we have a coordinate space; this is my Y axis; this is my x axis; so this is X and the Y, and this is my Center, so (0, 0). And what I'm doing here is simply I want to have Vector a; let's say this is just, um, Vector a, simple one with the coordinates, um, let's say four and -2, and I got Vector B—let me use a different color—Vector B that has coordinates, let's say -4 and 4. So let's actually visualize them; let's first start with the Vector a, uh, which has an x value of four—three, four—one, two, three, and four—and the Y value -2; so this is my Vector a. And let's now visualize the vector B, so -4 and 4, which means that—let me actually extend this—this is -4, so the x coordinate is -4, so it should be here, and then the y coordinate is four, so one, two, three, and four; it's this one, which means that my Vector B is this one. All right. So you can see now that the vector a is in here and the vector B is in here. Now what I want to do is to add these two vectors to each other. So what I want to do is to take this Vector a and add to this the vector B, which is: 4 + (-4) = 0, and then -2 + 4 is = 2; so (0, 2); it is zero and then two, two; so this is my result Vector.

So now when we are clear on how we can, in vectors, how we can perform these different operations and what it means in practice when it comes to looking at the vectors in a coordinate space and adding them or subtracting them, we are ready to look into the properties of vector additions. This is something that will definitely seem familiar to you, uh, from pre-algebra, where we are basically using all these properties that we already know that holds for, uh, numeric values, for the scalars, that being transferred to this Vector space. So we are going to talk about these four different properties that vectors have. The first one is the commutative property, which says that if we add a vector a to Vector B, then this is the same as adding a vector B to Vector a. So basically, the order of the vectors doesn't really matter when it comes down to adding them. So formally, A + B is equal to B + a for any vectors A and B of the same size. Then we have the associative property, which says A + (B + C) is equal to (A + B) + C; we can write both as A + B + C. Now what does this mean? We know from pre-algebra that this parenthesis means first do this addition and then do the rest of operations in here. It basically says if you add a to the B first and then you add the C, it's the same as first you add B to the C and then on the top of that you add the Vector a. So then the third property is addition of zero vectors, which says if we add a zero Vector to Vector a, then this is equal to adding a vector zero to a, and this is equal to Vector a. So adding the zero Vector has basically no impact on the vector whatsoever. Then the final property is subtracting a vector from itself, which means if we take the vector, we subtract the same Vector from itself, so a - a, and we get a zero Vector; so a - a is equal to the zero vector, and this yields the zero Vector.

Now let's look into each of those properties one by one, and let's, uh, look into specific examples; in some cases we will prove this on the example that we have to make these concepts much more clear. So let's start with this commutative property of vector additions. So we want to see whether A + B is equal to B + a. So let's say we have a vector a that has coordinates or magnitude and direction that is equal to (1, 2). Then we have a vector, uh, let's say B that has a magnitude and direction of (-2, 3). So the first thing that we want to check is indeed whether the A + B is equal to B + a. So therefore, let's first calculate this part, and then we will calculate this part, that I will define by one and two, and we will see whether we are indeed having the same value, the same vector, or not. So let's see. So we have here a, so A + B, which is the first value that we want to calculate: a + b is = to (1, 2) + (-2, 3), and we learned before that this is simply equal to: take this value and then add this one, so 1 + (-2), and then 2 + 3. So this gives us a vector: 1 + (-2) is = to -1, and 2 + 3 is = 5; so we get that A + B is = to (-1, 5), this Vector. Now let's look at the second quantity: so B Vector, B plus Vector a; this is equal to (-2, 3) + (1, 2), and this is equal to -2 + 1, and then 3 + 2. This gives us: -2 + 1 is = to -1, and 3 + 2 is equal to 5. So we can already see from here that the quantity one is indeed equal to quantity two, which proves that indeed the A + B is equal to B + a. What this basically means is that adding two different vectors, the direction or the order is not important; whether you add a on the top of the b or B to a, it doesn't matter; at the end it's the same. And actually you can also see it if you, uh, combine this or if you do this in more general terms. So let's say if we have a vector a which is equal to, in an N-dimensional space, (A1, A2 up to An), so it has n by 1 dimension, and you have a vector B with the same size from the same Rn Dimension, and it has element (B1, B2 up to Bn), and the dimension is equal to n by 1, then if we calculate first A + B, and this is equal to simply (A1 + B1, A2 + B2 up to An + Bn), and if you calculate the second, uh, amount, which is B + a, this is equal to (B1 + A1, B2 + A2 up to Bn + An), you can see that A1 + B1 is equal to B1 + A1, simply from pre-algebra you know that if those are all constants, for instance 2 + 3 is equal to 3 + 2; in the same way, A2 + B2 is = to B2 + A2, and then here up to An + Bn is equal to Bn + An. What this means is that all these elements they are basically the same, which means that we already have a proof. So we get this proof, and we can see that even for the general term, independent what this Vector a is, what this Vector B is, that a + b is equal to B + a. This is exactly what we saw before in the first property, which is called the commutative property of the vectors, that a + b is equal to B + a. Now let's move on to the other property, which is called the associative property of the vectors. Now what this property does and says is that A + (B + C) is equal to (A + B) + C, and this is then equal to A + B + C. Now let's then see, um, this specific property on an actual example. So what this basically says is that if we have this example where a is equal to—actually I had this before; let me simply just remove this part—let's then add our third Vector, which is C, and let's call it—let's say it has a representation of (4, 5)—then the idea behind this property is that what we need to prove here that A + B within the parenthesis + C is equal to A + (B + C), and then this is equal to A + B + C. So let's see actually whether this is indeed true for this specific case now.

This should come very intuitively, so I'm going to do it very quickly. So first, we have this quantity, this one; then we have this one; and the third one. Let's do it very quickly. So A + B + C is equal to one, two plus, and then we had C, so it is simply 4, 5. And then this is equal to—we saw before when doing this that we were getting 1 - 2, 2 + 3—and then we add this 4, 5. This is simply equal to 1 - 2 is -1, and 2 + 3 is 5 + 4, 5. Now, given that it doesn't really matter, no longer that we have uh, here, parenthesis or not, this basically means that this value is simply equal to -1 + 4. So here -1 + 4, here 5 + 5. So this is then equal to 3 and then 10.

All right, let's then now quickly do the second amount, which says: first add the vector B to vector C, and only then add the vector A on the top. What this means is that we need to take 1, 2—this is vector A—and we will only add this once we have added -2, 3—the vector B—plus to the vector 4, 5. Okay, so we can see that we are just leaving this in here. Let's first add this: 2 - 2 + 4, 3 + 5. So this gives us 1, 2 + -2 + 4 is uh, 2, and then 3 + 5 is 8. So this gives us—let me remove this calculation—so this gives us 1 + 2 = 3, and then 2 + 8 = 10. Okay, great. So now we got already the quantity 1 being equal to quantity 2. Let's check whether this is all equal to this one. It should already be um, something that you see now.

Given that um, we know just from mathematics that parentheses doesn't really matter when it comes to the scalers, and adding two vectors is basically very close to this idea of additive property um, of the additive property of the scalers, but just let's quickly do it to be 100% sure. So when we take this vector A to the B and to the C, we had all this. This is equal to 1, 2 added to -2, 3, and then added this to 4 and 5. Now, what this is equal to—let me actually write this in a bit shorter way, such that it can be all fit in in the small place—so 1, 2 + -2, 3 + 4, 5. This is equal to basically 1 - 2 + + 4, and then 2 + 3 + 5. Now, what is this number? 1 - 2 + 4 is simply equal to 1 - 2 = -1, and then + 4 = 3. So the first element is 3. 2 + 3 + 5 = 5 + 5, which is equal to 10. Perfect. So now we get the confirmation that indeed A + B + C = A + B + C = A + B + C.

So let's quickly also look into this addition of zero vector and the subtracting a vector from itself properties, and uh, the detailed explanation of this or example of this I will leave it to you. So when it comes to this A + um, 0 is equal to 0 + A is equal to A. So this property—let's say if A is equal to this 2, 3—and then we are adding on this A plus some zero vector, which basically means take 2, 3 and then added the same size of zero vector, you can see that this is the same as adding this zeros on these values. Now, what do we get? We get that this is equal to 2 + 0 is 2, and then 3 + 0 is 3. There we go. So we already see very quickly that it doesn't really matter whether we add a zero vector to this original A vector or not; we in all cases it just adding a zero vector has no effect. And seeing from the commutative property that A + B = B + A, we already know that if um, A + 0 = uh, A and is equal to this, then also 0 + A will be the same, and we can see indeed that we just saw that A + 0 is simply equal to A. So we basically have quickly proven all this.

Now, when it comes to the subtracting vector from itself, I think this is a very nice one just to see how we um, uh, take the same vector and subtract from that value and we get zero. And this is very similar to working with just real numbers, in the same way as 3 - 3 = 0. Also, when we have a vector consisting of the scalers, like A = 2, 3, in the same manner if we take this A and we subtract it from itself, so A - A, then what we will get is 2, 3 - 2, 3, and this will give us 2 - 2 is 0, and then 3 - 3 is 0. So we get a vector 0, so zero vector. So now when we are clear on how we can perform different operations on our vectors, and also we know uh, what are the properties of uh, adding and subtracting uh, different vectors, we are ready to move on to a bit more advanced topics. So uh, in this module we are going to discuss this idea of scalar multiplication; we're going to look into the example how uh, what happens and how we can do the uh, vector multiplication with the scalar; then we are going to uh, look into the span of vectors, what it means to have a span of vectors, uh, what is this idea of linear combination and the relationship between the span and linear combination and the unit vectors; then we are going to look into the application of scalar vector multiplication in audio scaling uh, example; and then finally we are going to finish off this module by looking into the length of a vector and a dot product, and we are going to uh, go back to this idea of distance, understanding vector magnitude and understanding vector L. So let's get started.

Now, before we look into this idea of span and linear combination, I quickly wanted to look into this idea of scalar multiplication and the um, specific definition of it. So formally, the scalar multiplication involves multiplying each component of a vector by a scalar value, effectively scaling the vector's magnitude. So what do I mean here? Let's say we have a vector, and I will write it in the general terms to keep everything general. So let's see we have a vector A, let me pick up my pen, A, and this vector A is from n-dimensional space, so it is from R<sup>n</sup>, and it can be represented by A<sub>1</sub>, A<sub>2</sub> up to A<sub>n</sub>, and I have this magnitude um, of a vector, and now I want to scale this uh, vector for which I know the magnitude and the direction; I want to scale it with a scalar, and we learned before that the scalar is just a number. So um, scalar in this case I will be uh, referring it to uh, by C, so C will be my scalar, and uh, this comes from R, which means that it's a real number. Let me actually use a different color to make it easier to follow. Okay, so my scalar will be with the color uh, red, so C, and C comes from R. So what do I mean by scalar multiplication? I mean that I want to find what is this C * A. This is what we mean by scalar multiplying with vector. Now, what does this definition say? It says when we are multiplying scalar we vector, so the scalar multiplication—meaning multiplying vector with the scalar—it involves multiplying each component of a vector by a scalar volume. So if we translate it to this specific example, it means that this amount, so this amount is equal to taking C and multiply find it with each element of this vector, so each component of vector. And what are the components of my vector? The A<sub>1</sub>, A<sub>2</sub>, A up to the point of A<sub>n</sub>, so all these components. So that means that the first element of this new vector, the scalar multiplication result, will be C * A<sub>1</sub>, C * A<sub>2</sub>, dot dot dot, so all this middle elements, and at the end again C times and then A<sub>n</sub>, and then in both cases of course the number of elements doesn't change, so the so the number of rows of my vector doesn't change; it's n, so here also n, and then the number of columns is the same, so it's just a column vector, so one column. So what we see here is that we go from A<sub>1</sub> to C * A<sub>1</sub>, we go from A<sub>2</sub> to C * A<sub>2</sub>, up to the A<sub>n</sub> transforms into C * A<sub>n</sub>. So we see very easily that I keep all the elements from this vector; I take them in here, and instead what I'm doing is that I'm multiplying every element from this vector by the scalar C. So this is exactly what this definition says. And let's actually go ahead and do a hands-on example with some real numbers to have this um, method and to have this uh, definition very clear in our mind, because we are going to make use of this fundamental operation scalar multiplication on and on in the upcoming lectures and just in general in your journey in any applied sciences.

So this is an example of scalar multiplication. Uh, here what we are doing is that we want to multiply this vector C, so in this case the vector is defined by a letter C, and then on the top we can see the arrow indicating that this is the vector now, and here we refer the scalar by a letter K; we are saying we want to perform scalar multiplication, which means that we want to multiply the uh, a vector C by the scalar K. So how we can do that? So what we want is to multiply K by C, and we just learned that for that what we need to do—let me write this over—so this = -2 multiplied by 4, -3. This is my vector, so this is the K and this is the C. This is equal to—so I take my scalar and I multiply it with the each of the element of the C—so -2 * 4 and then -2 * -3. So -2 * 4 = -8, and then -2 * -3, so minus it goes away, it becomes a plus, and 2 * 3 is 6. So my end result, the K * C is equal to -8, 6. This is my final result. So let's quickly also do yet another example, and this one is a unique one because it's relating to this idea of U multiplying something with a zero uh, which is something that we also uh, know from a high school that when we multiply number, let's say seven by zero, we are getting zero. And here in this example the uh, problem is: describe the effect of a scalar multiplication by zero on any vector, which means what we are doing is that in this example is we want to know what is this result of any vector, let's say vector uh, C. So we will use the same example C, only this time instead of multiplying it with scalar K = -2, our scalar will be zero, which means that C = 4, -3, and then K is now equal to zero, and we want to find out what is this K * C. Let me actually write down the K with a different color; K = 0. So what we want to find out is K * and then C, and this is that equal to 0. So I'm taking the K, 0, times, then I'm taking each of the elements of C, which is 4 and then -3, and I know that when multiplying the number with is 0 it gives me 0, which means that I end up with 0 here. 0 * 4 is 0, 0 * -3 is also 0, so I end up with a zero vector. Now this gives me an idea already that I can make a general conclusion that independent of the type of vector that I have, independent and what are this values in my C uh, if I have any vector C and I'm multiplying it with zero, then this will always give me a vector of zero, because all the members of this final vector will be just zeros. So if, for instance, the C comes from uh, let's say R<sup>n</sup>, so it has n different elements, it comes from n-dimensional space, then my final result of 0 * C, so this zero vector, this one, so zero, that this one will come also from R<sup>n</sup>, so you will be having a vector, so 0 * C will then be equal to 0, blah blah blah blah 0, so n times zeros. So this is then the idea of multiplying, so scaling a vector with zero, and this is our example two.

All right, so let's now move on on to our application of scalar vector multiplication, and then after this we will go back to this idea of linear combinations and spans. In this specific application we have a scalar vector multiplication, and we are looking into application of audio scaling. So the scalar vector multiplication audio processing uh, this can change the volume, for instance, of an audio signal without altering its content. So um, you might have noticed that um, when uh, when you are listening to video you can simply increase the volume of that video or decrease it, but you will notice that the content doesn't change; you are just increasing the volume or decreasing it. Even on the TV when you are watching a show you are increasing the voice or decreasing. Now what you're basically doing behind—and this is super interesting—is that behind the scenes what is happening is that there is simply um, audio that um, contains that show, and the audio of that show is being multiplied with a scalar, and that scale is simply the volume scale. If you scale it in such way that you want to decrease the volume, so the audio will then have a lower volume, then you are simply multiplying it uh, your vector containing the audio information in such way that those newer volume indications, they will be they will be containing lower numbers. Hope this makes sense. Let's look into the example; this make uh, this will definitely clear this out. So um, let's assume we have an a vector A that represents the audio signal, and we want to multiply vector A, A by scalar B to adjust the volume. So B is some sort of number; it can be, so B comes from R, so is a real number, while A is simply a vector. Given that it doesn't mentioning here, I'm assuming that A comes from R<sup>n</sup>, so it comes from R n-dimensional space. So imagine of A as this vector A<sub>1</sub>, A<sub>2</sub>, blah blah blah blah to A<sub>n</sub>, and each of these values it basically describes uh, an uh, the audio signal, so it represents um, uh, an amount, so it contains an amount that represents the audio signal of your uh, video or uh, your uh, show, and then the B in this case, for instance in this example you can see that the B is then uh, equal to for instance 1.2, 1/2, or B = -1/2. So you can see that B = 1/2, which basically is a fancy of saying that B = 0.5, or B can be equal to -1/2, which is -0.5. Now then it says then the B * A, which basically means multiplying our um, scalar beta by the vector containing the audio signal A, so this B * A is perceived as the same audio signal but at the lower volume. Now why lower? Because you can see that B = 0.5 or -0.5; it means that once you take all these elements of your A and you multiply it with a number that is smaller than one—in this case 0.5—then all these numbers will decrease, which means that also your audio volume will decrease. So let me actually uh, show you an example. So let's say our talk show is very short, and you know the audio variation is very low; you have a vector A that is quite small; it comes from a three-dimensional space, so R<sup>3</sup>, and it has numbers like 3, uh, 6, and then 5, so 3x1 vector, and then we have our audio adjustment scalar beta which is equal to 0.5. Now when we take the beta we're multiply it by our audio signal, then what we do times is clear, so times what we are doing is that we are simply taking all the elements of our A, so 3, 6, and 5, and what we are doing is that we are multiplying it by 0.5, 0.5, and 0.5, or you can also say 1/2. So what this is equal is that 3 * 0.5 is 1.5, 6 * 0.5 is 3, and then 5 * 0.5 is 2.5. And you can see that all this numbers 1.5, 3, and 2.5, they are smaller, and specifically two times times less than all the original values in the um, original audio. So original audio is A, which was 3, 6, and 5, and the new audio, the the scaled one is, so audio scaled, so B * A = 1.5, 3, and 2.5. So you can clearly see this transformation where this element 3 is larger than 1.5, 6 is larger than 3, and then the last element 5 is larger than 2.5, which means that this audio audio is much at a higher volume, so the volume two times higher than this audio. So this is basically the idea of uh, applying scalar multiplication to our audio pre-processing. I will leave the other example to you that will show that when your scalar is equal to -0.5, you again will end up with the lower volume, only that time the volume will be much much lower than the original one.

So now that we know how we can perform scale multiplication in theory, as well as we have looked into an example how we can do it in terms of the numbers and multiplying them, and we have also seen uh, applying scalar multiplication in practice uh, so we have seen in this audio processing stage the uh, multiplication process, we are ready to look into the visualization of it. This will help us to get a better understanding on uh, what exactly happens when we are scaling different vectors. Let's look actually in the following example. So let's assume we have a vector—oh, let me remove that—so let's assume we have a vector, and the vector is—let me get a color, this one, for instance, a vector A, and this vector A consists of elements 1 and 2. So where does this vector lie? The vector is with um, 1, so here in our coord system, this is our x-axis, this is our y-axis, and here we got uh, let me actually pick another color, let's say black one, and then we got 1 and then 2, right? This is 2, this is 1, so it is this one. So the line that we get here it is this one. So this is our vector A. Now let's assume I want to multiply my vector A, so I want to scale my vector A by a constant 3, by scalar 3. So I have a scalar, let's say I call K, and this K a different number, let's say K = 3. So what I wanted to do is to perform a scale of multiplication, so I want to obtain K multiplied by A, and we learned that this is simply equal to 3 times and then 1, 2, and then this is equal to 3 * 1, 3 * 2, which is equal to 3 and then 6. So let's also visualize this scaled uh, vector. So let me pick this yellow color; this will be our scaled vector. So we have done scale multiplication, and we are going to visualize that. So we have 3 and 6, so this is 3, 1, 2, 3, and this is 6. So we have this point. So you should already see what is going on here. So we got 3A here. So you can see that this part is our vector A, and this longer one is 3A, and even visually you can see that this longer vector is simply the three times of the shorter vector. So we got this, and then if you add on the top of this the same three times, you will then end up with the original, so scaled version of that. So basically this is A, this is A, this is A; we combine three different, so we scale A three times, and we simply get a three times longer version with the same direction. So you can see that when we are scaling, even visually it makes sense, so we are scaling our vector A three times, and we are

Just getting that vector, so we are transforming. Oh, let me remove this. So basically, we are taking this vector and we are scaling it up to this point. If I would do it only two times, then it would be something like this. Or one and a half times, it would be something like this, so only half of it. So now this should make much more sense.

Let us actually do yet another example to uh make sure that we are clear on this. Visualizations because we are going to make use of it when uh looking into this idea of linear combination in a span. So let's say we have a vector B, and this vector B has elements zero and three. So let's visualize and uh plot this vector. So it contains elements zero and three, so zero and three. So this is the X element and the Y element. On the Y axis, we can see this is three, which means that our vector B is this vector. All right, perfect. So this is our B. Let's now multiply, so scale our vector B by scalar tQ. So let's say we want to get 2 * B. So what is this amount? This is equal to 2 times, and I'm simply taking each of those elements, zero and then three. So this is then equal to 2 * 0 is 0, and then 2 * 3 is equal to 6. So this is my new scaled vector, 2 * times B vector, this one. So let's visualize this. The x-axis value is zero, so we are still here, and then the y-axis value is six. So what is six? This thing. All right. So you already should see that this is very similar what we had before. So this is 2B. All right. So this all uh should make sense. Uh also, we learned as part of the um high school when visualizing different plots. So this is quite similar to this idea of having Y is equal to X and then scaling it, getting like Y is equal to 2x. So in this case only we know exactly where the vector starts and ends. Uh so we have a much more specific definition instead of having all this infinite number of points on the line, but the idea stays the same. So we are taking this vector and we are then scaling it two times, so we get 2B vector. And I could do the same, only instead, what I could also do is I could do like uh 0.5 or 1/2 * B. So I take the half of it, which means I would get this vector. Or I could multiply it with minus one, so minus -1 * B. So I was scale with minus one, and and then I will simply get the negative version of my original vector, so this thing. This would be -B or -1 * B. So this is basically the idea of uh scaling multiplication when visualizing it in our coordinate system, Cartesian coordinate system. And now when we know all this, we are ready to move on on this idea of linear combin. And now when we know all this, we are ready to move on on this idea of linear combination.

So let's now formally define this ideal: linear combinations. A linear combination of vectors A1 up to Am using scalars B1 up to BM, or what we also refer as Beta 1 to Beta m, is the vector Beta 1 * A1 plus up to Beta M * Am. And the scalars are called the coefficient of linear combination. And any vector B in N dimensions can be expressed as a linear combination of the standard unit vectors E1 up to n. The coefficients in this combination are then the entries of B itself. Well, this is a whole bunch of information. Uh let's unpack them one by one. Firstly, um I want to mention about this m. So far we have seen this idea of n, so I just wanted P to experiment with a different one just to ensure that we are clear that you can use any source of identifier to describe the size of your um uh number of vectors that you got. And uh in this case we got M different vectors because so far we were using this n in order to describe the size of a vector. And now we are no longer talking about the size of a vector but the number of vectors; therefore, I specifically didn't use use the letter N. So here m is simply the number of vectors. So don't confuse this with this thing where we were plotting this and we were saying this A1, A2 up to An, because in here we basically mean that we are dealing with some vector a, and this has n different elements, whereas in here we are already moving from this idea of one vector, and now we are talking about m, multiple vectors. So we have M different vectors; they all look like kind of this, only with bit more complex indexing that we also saw before. All right, but we will learn this. Um that's not an issue. I just wanted to mention this to ensure we are at the same page. So then let's move on to this idea of using scalars Beta 1 till Beta M. So it's a common uh practice in linear algebra, in just in general in mathematics, but also definitely in data science, statistics, and in artificial intelligence to use Beta 1 as a way to describe the coefficient. So what do you mean by coefficient? It is just a scalar, so it's just a constant or a number. So in this case, for instance, this Beta 1 can be 0.5, Beta 1 can be uh two, Beta 1 can be let's say 100. It just describes how much we are multiplying, scaling this vector A1. So so far we have done a lot of scal and multiplication already, lot of details there, and we have seen different times different scalars that we use. We can use um zero as a scalar; we can use any other number as long as it's a real number. So this Beta 1 should belong uh in the a real number space, so it's a real number. And of course, the same holds for uh all the betas. So we have M different vectors, which means we are going to have M different scalars because each of those vectors we are going to multiply with their corresponding or respective scalars. So Beta one is basically the scalar uh or the um uh coefficient that we are using to scale A1. Maybe I can actually write this down on a new page such that we can save this as a SL Light page for you. Let's write it down. So what do we have as this idea of linear combination? So a linear combination simply involves taking several vectors uh to go from this uh formal definition to more practical uh terms. So we got this A1, A2 up to A, and what we want to do is to take the linear combination of this m different vectors. So we got m is the number of vectors, and to get a linear combination we need to uh scale each of those vectors, which means that we need to have this different scalars. Let's say Beta 1 for A1 and then plus Beta 2 for A2. So each time we are scaling each of those vectors where Beta 1 is the uh scalar or the coefficient of the vector A1, and we are multiplying, we are performing scalar multiplication of our scalar Beta 1 with the vector A1, and then we are adding to this our Beta 2, which is the coefficient corresponding to the vector A2, and then adding Beta 3 * A3 and then dot dot dot, so all these different uh vectors up to the point of Beta M times Am. And all this, so A1, A2 up to A, those are all vectors belonging to the m space, so those are all vectors coming from the um M dimensional space. So um in here this is the linear combination of our M different vectors, and the uh Beta 1, Beta 2 up to Beta M, those are all constants, so those are scalars or real numbers that belong to R, so those are real numbers. All right. So now when we are clear on that, let's also unpack this idea of coefficients. So the scalars are called the coefficients of linear combination. So basically all this members, so Beta 1, Beta 2 of two Beta M that belong to real number space, they are called coefficients. This is what we are referring as coefficients, and this coefficient, this IDE and name is super important because you will see this time and time again appearing in your uh very basic machine learning models or some other applications of linear algebra because the end goal is always to find these coefficients. So these coefficients, those are numbers that we are using to scale these different vectors, and uh the idea of coefficients is very central because those are numbers that define how exactly we are combining all these different vectors because this Beta 1, Beta 2, Beta 3, they can be different numbers, real numbers, and every time when we are choosing these coefficients or these betas, we will then end up with a different combination of these vectors. So we are basically mixing all these different vectors, and the way we mix it and how we will mix it it will depend on the values of this Beta 1, Beta 2, Beta 3 up to Beta m, so these coefficients. Therefore, coefficients are super important and they define the end results from our linear combination. So any vector B in N dimensions can be expressed as a linear combination of standard unit vectors E1 up to n. So when looking into this um idea of unit vectors, uh we saw already what this E1 is, what is E2 is up to En, and we saw that E1 is, for instance, if it's from an N dimensional space and it says from n dimensions, then E1 simply means 1, 0, 0, dot dot dot dot 0. Then E2 means 0, 1, 0, dot dot dot dot 0. So we already saw this; this is not something new that we are seeing, so 0, 0, blah blah blah, and then one at the end. And what this definition basically says is that any vector b, as long as the B comes from n dimensional space, we can represent this by using this uh unit vectors and by linearly combining them. So this is yet another part of this definition, and we are going to, by the way, um go through each of the parts of this definition one by one, going to each of the examples as well as visualizing them. So now I just want to quickly unpack all the parts in this definition before moving on to step-by-step examples and explanation. So this is about this linear combination of any n dimensional uh vector B that we can uh create by using a linear combination of these unit vectors. I will come to this in a bit. So then the final part of this definition is that the coefficient in this combination are the entries of B itself. So it says that the coefficients, so Beta 1 up to Beta m in this linear combination that we can create are the entries of B itself. So we will come to this section once we are done with the first part. So first let's have a good understanding of what this linear combination is and also touch base, and we will also formally define the idea of span, and after that we will move on on uh representing and expressing any vector B in N dimension space by using standard unit vectors E1 up to En and this idea of coefficients and then entries of B.

So let's start with the first one. So let's assume we have two different vectors. We have vector a, and this vector a is equal to 1, 2. So let's plot this 1 and 2 in our coordinate space, that is this one, which means that our vector a is this one. And let's assume that we have a vector B, and this vector B is equal to 0, 3. So 0 is here, and then 3 is here, which means that our vector B is this one. This is vector B. Now I want to create a linear combination of this vector a and vector B. So we just learned from the formal definition that in order to do so, I need a Beta 1 to multiply the vector a, and then I need Beta 2, which is the coefficient corresponding to to my second vector, in order to multiply the second vector, which is B vector B. Okay, so I'm getting the linear combination of A and B by taking any Beta 1 and Beta 2 which are real numbers. So Beta 1 and Beta 2 belong to R, so they are real numbers, and then I'm getting a linear combination of the two. So let's look into a few examples of a linear combination of vector A and B depending on the different choice of the coefficients like Beta 1 and Beta 2. So example one is that Beta 1 is equal to zero and then Beta Beta 2 is equal to 0. Now what is the linear combination of A and B when my coefficients Beta 1 and Beta 2 both are zero? It just means that I'm getting 0 times A plus 0 times B, which is of course 0 * 1, 0 * 2 plus and then multiplying vector B with a scalar zero, which is 0 * 0, 0 * 3. So let's quickly do this what this value is. This is equal to 0 * 1 is 0, 0 * 2 is 0, 0 * 0 is equal to 0, 0 * 3 is equal to 0, and this is then equal to 0 + 0, 0, 0 + 0 is 0. So I'm basically getting a vector zero. All right, so this is then equal to zero. So this equal to vector is zero. So I can also say that this vector or it's actually a point, so this point is simply a linear combination of these two vectors. Now this is a super basic case. Let's look at another case when our in our second example the Beta 1 and Beta 2, so our coefficients, they are actually not zero; there are some other nonzero real numbers. So in this example I will then take Beta 1 = to 3 and then Beta 2 is = to -2. And then what I will do is that I will take, actually I will take the um -2. Then I can also get rid of one of the elements, and I can get actually a zero for one of the elements. I will show you in a bit. So then the linear combination of A and B using these coefficients Beta 1 and Beta 2, where Beta 1 is equal to 3 and Beta 2 is equal to -2, is then equal to, so this amount, this amount is equal = 2 3 * (1, 2) and then plus we got -2 * (0, 3). Now what does this give us? 3 * 1 is = 3, 3 * 2 is = 6, plus and then -2 * 0 is = to 0, and then -2 * 3 is = -6. So you might have already noticed why I picked the Beta 2 equal to -2. I wanted these two numbers to actually cancel each other. So you see because 6 + -6 is equal to 0. So what do I get in my final result as a linear combination of these two vectors? I get 3 + 0, so 3 + 0, so 3 + 0 and then 6 + -6, and this gives me 3 + 0 is 3, 6 + -6 is 0. There we go. So this is my linear combination of the vector A and B when using the coefficients equal to 3 and -2 respectively. So this value is actually = to 3 and 0 in this case. All right. So let me actually clean this up because I also want to visualize this idea, and then we will go uh back to this uh linear combination. Let just summarize uh what we got before moving on to the plotting part. So if we simply take A and we add to this B, so this is the first case, so this is as you might have already guessed, this is also linear combination. Here we are saying take 1 * a and take a 1 * B, and this is yet in our linear combination. Here the Beta 1 is equal to 1 and then Beta 2 is equal to 1. So this linear combination gives us a vector that is 1 + 0 is = to 1, and then 2 + 3 is = 5. This is our first linear combination when the Beta 1 and Beta 2 is equal to one. This is a basic case, so doesn't require too much explanation. Here we have seen already this. Let's now look into the other example that we saw when we use uh the zeros as our coefficient, so that is 0 * A + 0 * B, then this gave us (0, 0). This was our second linear combination when Beta 1 and Beta 2 were both equal to zero. And then the third linear combination that we saw was that 3 * A + -2 * B. This gave us (3, 0). This was our third linear combination where Beta 1 was 3 and then Beta 2 was -2. So so then the linear combination of these two vectors is basically all the possible combinations of these two vectors that I can get when scaling or when multiplying these two different vectors by different sorts of uh vector by different sorts of scalars. So in all these different cases what I'm simply doing is I'm taking different sorts of coefficients Beta 1 and Beta 2 and then I'm getting the linear combination of these two vectors. We saw that in the simple case when we take A and we add to B, so basically the coefficients are equal to 1, so 1 * A + 1 * B, then the corresponding linear combination is equal to (1, 5). It means that we are getting this vectors, so (1, 5) is in here, which means that we are getting this one, this vector. If we get if we take the zero as a scalar, so Beta 1 and Beta 2 are both equal to zero, then the linear combination of these two vectors is simply the vector zero, which means that it is this point. Then if if we take the linear combination using 3 and -2 as coefficients, then we are getting this, so (0, 0) and (3, 0). So one, two, and three. This is three, then this is our linear combination. I can also take any other uh like scaled version of my B and of my A, and then I will get entirely different sort of vector. So let me actually show you a few more times um a couple of other examples. So let's say I keep my A, so I just take the Beta 1 equal to 1, but instead I scale my vector B two times. So this was at 3; I'm taking two times of my Beta, which means that I'm here. Then I can take this; I can add this to my A, so this is 2B. This will give me another linear combination of these two different vectors. I can also, you might recall that we said that the starting point and the end point doesn't really matter for for us. What matters is that we uh have the same magnitude and the same direction for our vectors. So this means that for me the vector being here and the vector being here doesn't matter. When I scale it with two, I can be here, with three I can be here. So this is the same as my B, only 3 * B. This is basically basically scaling B with three, and this in here means that my Beta 2 is simply equal to 3, and then this means that I can combine this with my A, which was in here, you remember. So this here, this means that I get yet another linear combination of these vectors, which means that I'm taking 3B and I'm adding this to my A, so 1 * A, my Beta 1 is equal to 1, my Beta 2 is equal to 3, which means that the linear combination of this is equal to (1, 2) plus and then 3 * B is equal to (0, 9). This is then the new linear combination, which is (1, 11). So the new linear combination is equal to (1, 11), so this thing, which is the same as this thing. And then you can go on and on. You can also calculate the same with a negative B, so you can take B and then you can scale it with -1, so this is -B, or you can go in here, in here. The same holds for A, so you can scale it all the way to here or in the negative side. So so this already uh give us the idea that we will go into to the next point, which is the span. So when it comes to the linear combination and in this specific case when we have these two vectors, we can combine these two vectors in anyway, and uh we can mix them up by using different sorts of coefficients of Beta 1 and

beta 2, and we will can we can get any Vector in our R2. So this means that any vector in our R2 we can represent by using only these two vectors, and this is not always the case. For this specific case, we are dealing with two vectors that we can use to represent any Vector in our r two. So what I mean here is that, let me clean this up, so independent what kind of vector you will give me in the R2, so it has two different elements, it is 2x one, I can use a linear combination of A and B. So a linear combination of A and B is beta 1 * a plus beta 2 * B in order to represent this Vector X1 and X2. Therefore, we are saying, and we will come to this um in the next slide too, that the spend of the vectors A and B, so this is the set of all possible combinations of these two vectors, is equal to R2 because any Vector in R2 can be represented as a linear combination of these two vectors. So we have a linear combination of A and B, and I'm saying that I can represent any Vector, so here vector X, and this Vector X I'm representing by X1 and X2, and X1 and X2 can be any real numbers. So X1 and X2, they belong to R, and X is simply a two-dimensional Vector. So X1 and X2, those can be any numbers: 0, 1, 2, 100, anything, and I'm saying any number in this two dimensional space. So whether it is this one, any Vector, this one, this one, or this vector, or this one, any Vector that you give me in two dimensional space, I can find a linear combination of this A and B that is equal to that Vector. So I can represent that Vector as a linear combination of vector A and B that we saw before.

Let's actually prove that. So I'm going to represent this uh X1 and X2 by a linear combination of this Vector A and B, and how we can do that. So we have beta 1 * a plus beta 2 * B, it is equal to X1 and X2, where beta 1 and beta 2, so beta 1 and beta 2, they are constants, so they are also real numbers. Let's unpack this, which is beta 1 * 1 2 + beta 2 * 0 3, and this should be equal to X1 and X2. Now this is equivalent of: so beta 1 * 1, beta 1 * 2 plus beta 2 * 0 and then beta 2 * 3, and this should be equal to X1 X2. So this is my beta 1 a, this is my Beta 2 B, and this is my X, all right. So now what we get is that, and this is equivalent, beta 1 * 1 is equal to beta 1, beta 1 * 2 is 2 beta 1, plus then here beta 2 * 0 is 0, and then beta 2 * 3 is 3 beta 2. So we have learned U from the uh operations on the vectors that beta one uh, so in this case when we are adding two vectors, so beta 1 + 0 is the uh amount that we need to put as our first element. So when we are adding two vectors we just need to take their corresponding elements, we need to add them up, so equal to beta 1 + 0, and then 2 beta 1 + 3 beta 2. This is the result, and this should be equal to X1 and X2, at least this is what I'm claiming. So this zero doesn't matter, so we what we are getting from here is that beta 1 is = to X1, and then 2 beta 1 + 3 b. 2 is equal to X2. This is the two expressions that we are getting based on all these different calculations. So let me actually remove all this. So we have beta 1 is equal to X1, and 2 beta 1 + 3 beta 2 is equal to X2. So here, given that we have already that beta 1 is equal to X1, and here we have two unknowns, I'm going to fill in the value for beta 1 in here. So I'm going to take this and I'm going to fill in it in here, so for this volue. So remember that beta 1 and beta 2 are two unknowns, and each one and next to are just uh numbers that we will get when we uh know exactly the vector and we just want to represent the vector as a linear combination of two vectors. So when I take this uh value for beta 1 which is equal to X1 and I'm going to fill that in in here, it means that I'm going to get from here that beta 1 is equal to X1 and 2 * X1, because beta 1 is equal to X1, and here I got beta 1, I'm just filling in that value for beta 1 which is equal to X1, so 2 * X1 and then the rest I'm just taking over 3 beta 2 is equal to X2. Let me remove this, and from here what we are getting is that beta 1 is equal to X1, and I will solve this equation for the unknown which is equal beta 2, so I will just take the three beta 2 from left hand side, I will leave it there, and I will take this and I will take it over to the right, so I'm taking X2 over and this two X1, so this part I'm just taking to the right two of the equation, so Min - 2 X1, which then on its turn is equal to, so it goes to B 1 is = to X1, and then beta 2 is = to X2 - 2 X1 / 2 three, perfect.

So what do we get here? What is our end result, and why is it significant? So what we are getting here is that, based on all this information, without knowing beta 1 and beta 2, we got that beta 1 should be equal to X1, and beta 2 should be equal to X2 - 2x1 / to three. This means that independent what kind of x's you will give me, so what kind of vector we have in our R2, so this X1 and X2 they are just real numbers, we can always find beta 1 and beta 2 that we can use to represent that X1 X2, so our X Vector as a linear combination of these two vectors. Let me actually give you an example. So let's remove this. So let's assume we have a vector X, and this x is is equal to four and let's say three. So if we got this Vector X, and we are saying we can use this Vector A and B to represent X as a linear combination of vector A and B, which means that I can find I can find real number beta 1 and beta 2 that I can use to multiply the vector A and B respectively, combine them together, so their linear combination that will be equal to this Vector X. So this is my X1, this is my X2, so this is equal to 4 and 3. Now using this, let's actually see whether that is true. So based on this example, my beta 1 should be equal to X1 which is four, my Beta 2 should be equal to X2 which is 3, so beta 2 should be equal to X2 which is 3 - 2 * X1 which is 4 / to three, and what's this number? This means that my beta 1 should be equal to 4 and my Beta 2 should be equal to 3 - 8, so 3 - 8 / to 3, and this is equal to Minus 5 / to 3. So this means that I use a coefficients beta 1 is equal to 4 and beta 2 = to - 5 / to 3 to represent my Vector X as a linear combination of vector a and Vector B. Let's actually prove that too as a final step. So let's see where the four times Vector a which is 1 2 + - 5 / to 3, whether this is indeed equal to Vector X. So my Vector B is 0 3, so this is the first part, and I want to prove that this is indeed equal to X, and we already know what x is, so this is equal to 4 * 1, 4 * 2 plus, and then here we got - 5 / 3 * 0 and then - 5 / to 3 * 3. This is equal to 4 * 1 is = to 4, 4 * 2 is = to 8, and then here we need to subtract minus 5/3 5 5 / 3 * 0 is equal to 0, so this one is zero, and then minus 5 / to 3, so 5/3 * 3, this ones are canceling out and we got 8 + - 5, so here the plus and here minus just to make sure we got everything right, and this is equal to 4 and then 8 + - 5 is equal to 3. So you can see already that this amount that we got here is equal to X which was equal to 4 / to3, and this helps us to uh verify and to know for sure that indeed while given any Vector in a two dimensional space in R2 X, independent what this X1 is or X2 is, we can always find a pair of beta 1 and beta 2 that will ensure that the beta 1 a plus beta 2 B is actually equal to this x, where X A and B they are part of R2 and a is equal to 1 2 and then B is equal to 0 3.

So we can represent any Vector in our two-dimensional space as a linear comp combination of this Vector a with elements 1 2 and um Vector B with elements 0 3, and that's exactly what we saw here because we could find any vector and we can represent this Vector as a linear combination of this Vector A and B, this Vector as a linear combination of this A and B, this Vector as a linear combination in any vector or a point in this plane we can represent as a linear combination of this Vector a and Vector B, and in this specific case with this Vector a and Vector B we are saying that Vector a and Vector B they spin R2, so Vector A and B span R 2. Now we will come to these definitions of the span and uh just in general for different sorts of vectors we will see what this IDE of span is, but for now given that we just proved that we can represent any Vector in R2 as a linear combination of these two vectors A and B, therefore we can say, and we usually say it in linear algebra, that the vector a and Vector B they spend R2. Before moving on onto this concept of Spence that we just touched upon in our example, I wanted to quickly go back to this example that I promised to discuss uh which was part of the definition of the linear combinations and unit vectors because we saw in our definition and let me just show you that uh the uh definition was providing these two highlights, these two bullet points, and was saying any Vector B in N Dimensions can be expressed as a linear combination of the standard unit vectors E1 to up to e n, and the coefficients in this combination are the entries of B itself.

Let's look into the example and see what we mean by that. In this specific example, we have this Vector B, it is coming from the three dimensional space which we can see given that we have three different uh three entries, so three uh elements in our Vector, so it's 3 by one, and this means that b belongs to R3, and in here we can see that B can be written as a linear combination of these three vectors. So you can see that b e is equal to -1 * this Vector 1 0 0, so this one, then we have + 3 * 0 1 0 Vector, so this one, and plus 5 * this third Vector which is 0 0 1. Now we already know from the unit vectors that E1 is equal to 1 0 0, assuming that we are in three dimensional space, E2 is equal to 0 1 0, and E3 is equal to 0 0 1. You can notice that that's exactly what we got here, this Vector is E1, this Vector is E2, and this Vector is E3, where E1 E2 and E3 belong to three-dimensional space, okay. So another thing that we can see is that here we got coefficients minus one, here three, and here five, so this is basically how beta 1, beta 2, and beta 3, using the common conventions that we saw before when describing the linear combination. So let's actually check that, and then we will comment on these values. So let's check whether -1 * E1 + 3 * E2 + 5 * E3 is indeed equal to this B, so this is equal to -1 * this Vector gives us -1 0 0, three times this E2 gives us 0 3 and zero, and then 5 * E3 gives us 0 0 5, and this is equal to -1 + 0 + 0 is = -1, 0 + 3 + 0 is = 3, and then 0 + 0 + 5 is equal to 5. Now what do we get here? We see that this which is equal to this, it is equal to this Vector B indeed. Okay, so now when we have indeed checked that B can be represented as a linear combination of this three vectors, this unit vectors E1 E2 E3, another thing that we can notice, and I'm sure you already did, is that those coefficients they are not just randomly picked coefficients, those are the entries of this Vector B. So this is exactly what that definition was about, it was saying that any Vector B, including this example in in this case threedimensional space, can be Express as a linear combination of the standard unit vectors E1 A2 E3 Etc. So this coefficients in this combination, so you can see that the beta 1 beta 2 and beta 3 which are our coefficients in our linear combination, they are the entries, so this values of the B itself. So the same will hold for four-dimensional case, five dimensional case, n dimensional case. So this means that if we write this down for General case, just to ensure that we are clear on this part of the definition, so if we got B Vector in N dimensional space, so it got B1 B2 up to BN as the elements of it, comes from RN, then we can represent this B as a linear combination of unit vectors coming from the N dimensional space, so we got E1 E2 up to e n that belong to n dimensional space, and we can represent this B as a linear combination of these unit vectors, so by using beta 1 time, so this this is a common Convention of the coefficient as you Rec called time C1 then B beta 2 * E2 blah blah blah plus beta n * e n, and what is important here is that this beta 1 beta 2 and beta n those are not just some coefficients, but we already know what these coefficients are because we can then represent this beta by taking the values, so those are the entries, the elements of the vector B itself, so it is B1 * B1 plus b2 time E2 dot dot dot plus BN times e n, where B1 B2 up Q BN they are all real numbers.

So basically, knowing what these Vector is, these elements of this Vector, we can always describe and express it as a linear combination of the standard unit vectors, and if you're wondering why is this important, in some cases when performing different operations or working on different algorithms it just becomes handy to represent your vector as a linear combination of multiple vectors, and in those cases exactly you can make use of this property of linear combinations to express your n dimensional Vector b as a linear combination of the standard unit vectors because everything is then down to you by having this Vector B you will already know what are the entries that you can use as your coefficients in this case beta 1 beta 2, so those are all these values coming from your vector itself, and then the remaining is also none because you know exactly what these unit vectors are and how they are represented. So here for instance the E1 is basically one 0 0 blah blah blah blah 0, and then this is n by1 Vector, here the E2 is equal to 0 1 Z blah blah blah and then zero here, so n * 1 again up to the point where you have the N where you have all the zeros only the last element is one again n by one vector. So this is the idea behind this second part of this definition which says that any Vector being in N dimensional space can be expressed as a linear combination of the standard unit vectors E1 up to e n. Let's now talk about other concept which is also super important which is the span of vectors. So by definition the span of a set of vectors is a set of all possible linear combinations of these vectors. So if V is equal to V1 V2 up to V K and is a set of vectors, then the span of V is written as a span V and it includes any vectors that can be expressed as C1 V1 up to C2 V2 up to CK VK. So basically it is a common uh notation uh to say that if we got for instance vectors V1 V2 up to VN, so we have n different vectors, then we say that the span of V1 V2 up to VN that this is simply the notation that we use in order to describe the span of these vectors, and we briefly spoke about this concept of span when we were looking into our example that we saw before. So you might recall vectors A and B that we had and we saw and we said that the span of a and b is the entire space in the two dimensional uh real number space, so we said that span of a and b is equal to R2 where our Vector a was simply equal to one 2 and B was equal to 0 3. So we proved that the span of one two and 0 3 was the entire R2 and how we knew that because we proved that any Vector in R2 could be represented as a linear combination of these uh two vectors. So you might recall that we solved this equations, we saw that in depend what kind of X1 and X2 uh one will give us we can always use the uh um we we found this amount, let me see where I can find it back, I no longer have this, so we saw that for uh specific values of um beta 1 and beta 2 we can always get a linear combination of this A and B in order to get our desired factor x, so beta 1 * a plus beta 2 * e will always then be equal to X1 and X2 if our Vector a and Vector B are those, but of course this doesn't hold for all the vectors, so not for all two-dimensional A and B uh we can say that the span of these vectors is the entire R2. Therefore, to better understand this concept of span and this concept of span of vectors I wanted to distinguish five different cases, one of which we already spoke about and that is the case when we had this Vector a and Vector B and we said that the span of a and b is the entire R2, but we will also look into the case when we for instance have a span of the zero Vector, the span of a single vector, and the span of perpendicular vectors. We might also look if there is time left we will also look into the span of parallel vectors. So let's now look into this cases one by one. So let's say we have a vector of zero, so we have a zero vector, so this is a very simple case, we will start with the simplest case and we will move on B to two more advanced cases. If we have a vector a that is a zero Vector 0 0, then independent what kind of scaler we will use to scale this, so let's say um we Define it by C, so C * Z, independent what kind of scale we will use this will always end up being equal to 0 0. So if C is equal to 0, C * 0 will be equal to 0, if C is equal to 1, C * 0 will be equal to Z, or C is equal to 100, C * 0 will still be 0 0. So independent what kind of scal we will be using, what kind of lead linear combination we will create from our Vector a, this will always stay in here, so the point the vector will always stay in here in our two Dimension space. So this is completely different from what we saw before when we could create and we could take any Vector in our R2 and we could represent it as a linear combination of these two vectors that we saw in the previous example. So in this specific case um scaling the zero

With independent of any scalers we use, this will not change the magnitude, nor will it change the direction of our vector. So, no matter how we scale it, we still get zero. This means that the span of the zero vector is just the zero vector itself. So you can see that independent of what I scale the zero vector by, I always end up with the same zero vector. So therefore, the span of the zero vector is equal to zero. Because by definition, the span of a set of vectors is the collection of all possible vectors that I can reach by performing linear combination, and in this case, I will always reach this zero vector. So all possible collections of these vectors are the vector 0, 0, which is a single vector and the same as the input. So this is the basic case. Now let's move on to a bit more, uh, advanced case; so, a bit more complicated than this one, but itself also very easy, which is when we got a single vector, a.

So let's say a is equal to one and two. Now I want to know what is the span of a. In order to know what is the span of a, we simply need to understand what are all these possible collections of vectors that I can get when I'm, uh, combining, um, a; I'm multiplying a with different coefficients. So what are all the possible linear combinations of this vector? Because I got just a single vector, a, so a is one, one, two, which means one in here and then two here; my a is this vector. And let's look into, uh, different, uh, scalar multiplications of this vector. So let's say I want to calculate c * a, so the scalar multiplication of this, where c is equal to 2, c is equal to 3, c is equal to -1, c is equal to -3, and of course, c is equal to 1. So in all these cases, when c is equal to 1, then the linear combination, in this case just the scalar multiplication of this single vector a, so 1 * a is simply equal to one and two; so the same vector a. c is equal to 2; this will give me 2 and 4. c is equal to 3; this will give me 3 and 6. c is equal to -1 will give me -1, -2 for my a, and then c is equal to -3 will give me -3 and then -6.

So let's plot each of those. So if we got, for instance, c is equal to one case, you can see that we already got that vector in here, so it is this vector. When c is equal to two, then we got this one, so two and four; so where is that? It is in here. Let me use another color; it is in here. When when we got c is equal to three, then we got, so this one, three and six; so this is three and this is six; so it gives me this vector in the next example. So in the next linear combination, we have c is equal to minus one, so we got -1 and -2; so where is -1? It is in here; where is -2? It is in here; so I'm getting this vector. And then finally, when I have, let me change the color, when I have c is equal to -3, so this case, then I got -3 and -6, which means that here is my -3, here is my -6, so we got this thing. So you already should see what is going on here. When we got just the single vector for which we need to know what is a linear combination, and that vector is not a zero vector; it has nonzero elements, but, um, it's still, it is just a single vector, then all its linear combinations, given that it is simply a scaled multiplication of it, we are all getting them on the same line. So you can see all the linear combinations of this single vector is just a scaled version of it, and it lies on the same line. So what this tells us is that essentially you can move along the line defined by this vector a, but you cannot leave it; so you cannot get a vector that is in here, that is in here, that is in here, in here; so you cannot leave this, uh, line; you will always stay on this line. So this line essentially is the span of a.

So when we got a single vector, and that vector is not equal to the zero vector, then the span of a is equal to, and this can be expressed as c * a, given that the c is a real number. So we already saw this; independent of what kind of scalar we will take, any linear combination of it will end up simply the c * a. So therefore, we are generalizing this, and we are seeing that the span of a, so the set of all possible linear combinations of this a, is simply equal to c * a, given that the c is a real number. This is basically the span of a real, uh, uh, vector in a two-dimensional space. Let's now look into the next case, the next example, when we will calculate or we will define the span of perpendicular vectors. So let's look into another example when we are looking for a case when the, um, when we want to find out the span of perpendicular vectors.

So imagine we have these two vectors, vector a and vector b, where a is equal to 1, 0, and then b is equal to 0, 1. So we are still in our lovely, uh, two-dimensional space. So let's first visualize the vector a; it's quite basic; it is this one. And then vector b, it is simply this one. So we can already see why they are perpendicular. So you can see that they are forming this, um, 90° angle; so a right angle here. And then we know that the span, so that's exactly what we want to find out, so the span of a and b, and this is what we want to find out, and we know that the span of two vectors is the set of all possible linear combinations of these vectors. So we want to see what are these all possible linear combinations of c1; so all the possible outcomes that we will get when we get linear combinations of these two vectors. So basically, c1 * a + c2 * b, because those are all the linear combinations of these two vectors; c1 * a + c2 * b. Nothing that we can see here quickly is that c1 * a, so this part, those are all the scaled versions of a; so scaling multiplications of a. And this second term in the linear combination, those are all the scalar variations; so scalar multiplications of vector b, which means that, and we already have seen this time and time, uh, again, that when it comes to vector a, all its linear combinations they will lie on the same line. So let me take this color; so if I do 2a, so c1 is equal to c2, then I will be in here; if c1 is equal to three, then I will be here; c1 is equal to 4, I will be here; c1 is equal to 10, I will be in here. And then the opposite holds as well; if c1 is equal to, for instance, -2, then I will be in here; if it's equal to -5, I will be, my vector will look like this, and so on. So this means that all the scaled multiplications of vector a will lie on this line. So I can also say that the span of, see, the span of a, so span of a is simply equal to c1a; so you can see in here, so on this line basically; so this is c1a. So this basically means independent of what kind of c1 I will take, whether it is 1, 2, 3, 0, -5, -100, I will always end up on this line; so this line. So this is about the, uh, scaled multiplication of a, but of course, to create this linear combination of a and b, we also have the second element, which is the all possible scaled multiplications with a vector b; so c2b2. So let's see what that looks like. So if I, for instance, take c2 is equal to 1, I will be in here; if I take c2 is equal to 2, I will be in here; c2 is equal to 5, I will be in here; c2 is equal to -5, I will be here. So you are already seeing what is happening here; so all the possible scaled multiplications with vector b will be on this line.

So now we are then getting that the span of b will then be equal to c2 and then b. And here I'm not, uh, using formal notation; I'm just trying to, I'm just trying to, uh, draft the idea of the span of vector a and, uh, span of vector b, because we are not, uh, done yet; we still need to combine the two in order to find the span of vectors a and b when they are perpendicular. So when this angle is simply 90°. Okay, so let's also add this on our plot; so this is c2 and then b. So this already gives us an idea that all the possible combinations of the two; so when we add these two elements to each other, the outcome will always lie on these two lines, but there is no way that we can find any other coefficient for c1 or c2 that can help us to get a value that will be, so a vector that will be in here or in here or in here or in here; that's just not possible. So just you can try to go ahead and solve that equations like we did before, before, and you will see that there, there is no way that you can pick here a line and you can represent it as a linear combination of these two vectors; it just not possible. And later on we will see why, but just keep in mind for now that once we have this, this type of vectors, when two vectors are perpendicular, then, um, we cannot find a line, a vector that is outside of the two lines. So here you can see the x-axis and the y-axis, but it can also be like this; it can also be like this, but then you cannot find any other line that lies outside of this area that you can, uh, create a linear combination of these two different vectors, and then you say, then you cannot say that you can create a linear combination of these two vectors a and b. So therefore, when it comes to defining the span of the two perpendicular lines, we say that the span of a and b, given that a and b are perpendicular, but also given that these values, in this case, you know, a is equal to 1, 0, b is equal to 0, 1, then their span, you might have already guessed, is equal to c1a + c2b, given that c1 and c2 are of course real numbers. So in this case, c1 and and c2, as expected, are just scalars; so they are just some real numbers coming from R, and, uh, this a and b, those are vectors that are being spanned, and in this case specifically the vector a is equal to this one, zero, and vector b is equal to 0, 1. And this expression that we see here, this span, this simply describes the set of all possible vectors that can be formed by adding the scaled versions of this a and b; so c1a + c2b in order to form this linear combination. So this set, this set of c1a + c2b, which is a linear combination, all possible linear combinations of the two vectors, so, um, it effectively covers the entire plane, illustrating that any point in 2D space can be reached by some combination of a and b.

Let's now move towards our final example that we saw also as part of our definition for the span of vectors, in order to check and to learn how we can usually check, uh, whether the two vectors they really span the entire space. So in this case we got two vectors; we got vector, uh, v1, which is equal to 1, 2, and vector v2, which is, it's equal to 3, 4. So in here, and also, um, we, uh, have in our example that it says the span of v1 and v2 is all over the R2, because any vector in R2 can be expressed as a linear combination of v1 and v2. So the example basically is saying that if we know that, um, we can express any vector in R2 as a linear combination of v1 and v2, then we say that the span of v1 and v2 is the entire R2. So let's actually go ahead and prove that from our example. So we have x, which we can represent as x1 and x2, and x1 and x2 are just real numbers, and we got v1, which is 1, 2, v2, which is 3, 4, and we got in our example that, uh, we need to prove that the span of v1 and v2 is the entire R2. So for that, what we need to do is we need to prove that we can express our coefficients c1 and c2 in such a way, using x1 and x2, that independent of what these x1 and x2 are, so what kind of x vector we have, whether this is like one, two, or this is 0, 4, or this is 0, 0, 0 and, uh, 5,000; independent what kind of vector we give here, so x1 and x2 values, as long as those are real, uh, numbers, we can always find a set of c1 and c2 that we can use as coefficients in order to create a linear combination from vectors v1 and v2, and in that case we say then the span of v1 and v2 is the entire R2. Okay, so let's go ahead and actually prove that using our previous knowledge that we already gained. So keeping in mind that c1 and c2 are unknown numbers for us, whereas x1 and x2 are just a way to describe those elements in our vector x that will be provided to us. So x1 and x2 will be basically known, and c1 and c2 are the unknowns that we are chasing. So for that, the first thing that I'm going to do is to describe this linear combination that we got here, c1v1 + c2v2, with actual equations, unknown equations. And the way I'm going to do it is by simply filling in this vector v1 and vector v2, um, values. So we have c1 and then c2 here, and then here I got 1, 2 plus, and then 3 and 4 here, and what is this amount? So this is equal to, let me actually go on to the next row so we can create, um, a set of equations for this. So this is equal to x, and we get c1 * 1 + c2 * 3 is equal to, and remember that this is x, and we said that the x is equal to x1 and x2, so this basically equal to, we can write here, is equal to x1 and x2. So this is then equal to, let's not skip all the steps, x1 and x2. So here then the second elements need to be added, so c1 * 2 and c2 * 4, which is then the same as c1 + 3c2 and then 2c1 + 4c2, and then this we are saying this vector is equal to x1 and x2. So this is what we have here. And let's move from the vectors to equations. So given that we have this, we are allowed to say that this gives us actually two equations; this means that this element, this element from this part should be equal to this, and this element should be equal to this. Now let's write it down; we see that c1 + 3c2 should be equal to x1; 2c1 + 4c2 is equal to x2. This is all that we see in here. Let's remove this to keep the space clean. Now what this means is that we have two equations with two unknowns, c1 and c2, and x1 and x2 are the numbers that will be provided to us as part of our vector. So what we want to prove is that we can describe and we can express c1 and c2, which are our unknowns, using x1 and x2. So you see here, this is c1, c2, those are our unknowns, and we want to describe them by using x1 and x2, and very soon we will also see why. So for now, let's try to express those two unknowns using our knowns like x1 and x2. So here I already see that c1 is alone, so there is no scalar, so I will make use of that opportunity to keep the c1 on the left hand side, and I will take this, this amount to the right, so I will say c1 is equal to x1 - 3c2. Two; slightly better. So I have c1 at the left; I do have x1 in the right, but I also have 3c2 in here. But another thing that you will notice is that in my second expression here, I got 2c1 + 4c2 + x2; I want to have the c2 only, because then I will have an expression of my c2 only using x1 and x2, um, so numbers that are that will be provided to me that are known. So for that, what I'm going to do is basically trying to solve, uh, two equations with two unknowns, exactly the same, um, process. So I'm going to take this c1 from the first equation, and I'm going to fill in in the second equation. So I am going to say two times, and here I'm going to fill in that c1 expression from here, so x1 - 3c2 / 2, so this is my c1, plus just taking over this part, so 4c2 is equal to x2. Okay, perfect. So now what I end up with is c1 is = to x1 - 3c2, just taking it over, and then here I'm opening parenthesis, which is 2x1 - 6c2 + 4c2 is equal to, here I forgot an x2, is equal to x2. Okay, one step closer, why? Because I in my second equation I no longer have a c1; I only have a c2, which is great, which means that this gives me an indication that I can rewrite the c2, which is unknown, with knowns, with x1 and x2. So let's make use of that opportunity; the first equation I would just take over, so c1 is equal to and then x1 - 3c2, and then here I will do 2x1, and then here we got 2 * c2s, which means we can combine them, so -6 * c2 + 4c2, it gives me -2 * c2, and this is equal to x2. All right, let's now solve that part. So what I want to have is just the c2 in the left hand side, which means I need to bring all this to the right, and I need to get rid of them such that I can leave the c2 in the left entirely alone. So this is what I'm basically chasing. For that, I'm going to, once again, rewrite c1 is equal to x1 - 3c2, and this time I'm going to take them -2c2 here; I'm going to leave that in the left, but then this one I'm going to bring to the right, so x2 - 2x1. Here I need to be very careful to not make a mistake that will mess up my entire calculation. All right, so now we are one step closer, just taking over the first equation again, so c1 is equal to x1 - 3c2, and here what I need to do to get rid of this minus two is to divide the two sides, so both / -2, that will help me to keep the c2 only in the left alone without any scalar. So the c2 is then equal to x2 - 2x1 / -2. So this is what I end up with. Perfect. So we are very close; stay with me. So, uh, here what we are getting is that c2 is equal to this amount; we see that now we no longer have any other c in here, which is great. And remember that x1 and x2 will be numbers that it will be provided to us; I just wanted to give everything general. And then, uh, another thing that I want to fix is this c2, because this c2 is an unknown, and I want to fill in, uh, this value of c2 in here such that for the c1 I will have a similar picture; so in the left hand side I will have c1, in the right hand side I will express the c1 with known numbers, so x1 and x2, but not the c2 or others. All right, so let's then go ahead and do that. First, I will write c2 in a simpler way, so c2 is equal to, here I got a minus, I will just write here minus, so I will take the minus over here, and then I will write x2 - 2x1.

To be super careful with this minus, therefore I'm using parenthesis. So now I'm going to use this C2, and I'm going to fill that in in here. So C1 is equal to X1 - 3 * I can also make it plus because minus of here, so minus of here and minus of here will cancel out. Therefore, I will do plus three times and then X2 - 3X1 / 22. And this then gives me C1 is equal to X1 + 3 / 2 * X2 - 3, sorry, 2, almost made a mistake, 2X1. And then C2 is = to X2 - 2X1 / 2 2. And here we got minus. Perfect. All right, awesome. So now we have expressed X1 and X2. Well, careful with this X1 and X2, we only noun numbers now. What I'm going to do is that I'm going to prove that independent what kind of X we will be taking here, we will end up getting the C1 and C2 using this what we just found here that will give us a linear combination of these two vectors that will be equal to that eight. So for that, so to prove that this pen of V1 and V2 is the entire R2, I need to prove that independent what kind of X I will take, so X1 and X2, I can always find the C1 and C2 that I um just calculate in here using that X1 and X2 that I can then use to combine with my V1 and V2 to find the linear combination of these two vectors with that C1 and C2, which will be equal to this x. So for that, what I need to do first is to take such a uh random X. So let's say my X is equal to (0, 4). This means that my X1 is equal to 0 and X2 is equal to 4. What this means is that this gives me C1, which is equal to and here X1, so I'm basically filling these two values for here to obtain my C1, so C1 corresponding to this specific Vector X. So X1 is equal to 0, which means I end up C1 is = 0 + 3 / 2 * X2 is = 4, so 4 - and then 2 * X1 is equal to -2 * 0, which is 0. And then C2 is = to - and then X2 is equal to 4, so 4 and then - 2 * XY is = to 0, 0. And then this divided to two. Now what are those numbers? So C1 is equal to 3 / 2 * 4, which is 3 * 2, so 6. And then C2 is equal to -4 and then minus, so this is zero, this cancels out, which means 4 / 2 is 2, and then C2 is equal to -2. So basically, I have calculated the coefficients C1 and C2 by just knowing what is this Vector. So knowing X, the provide X1 and X2, I have calculated my C1 and C2 using my derivations in here. So let's now get rid of this, this calculations, to clear some space and to do the final part, which is compute the linear combination of vector V1 and V2 for this specific coefficients. Well, knowing what this given Vector now is, example random Vector, so the C1 is equal to 6, which means 6 * and then Vector V1 is (1, 2). So this is first part of my linear combination plus and then C2 is equal to -2 times then here (3, 4). What is this? This is equal to (6, 6 * 2 is 12) plus now let's calculate the second part. -2 * 3 is -6 and -2 * 4 is -8. So what does this give us? (6 - 6) and (12 - 8). This gives us (0, 4). Nice. So this confirms that we have done everything also correctly, which is great because we have seen that using this C1 and C2 that we have just calculated, we have successfully uh computed the 6V1, so linear combination of this uh two vectors V1 plus -2, -2 and then V2, and we have seen that this linear combination is equal to (0, 4), which is exactly our x. So in this way we have proven that independent what kind of vector we will pick, what kind of X we will pick here, we can always find and calculate the corresponding coefficients C1 and C2 in the same way as I just did. And then by using those when we calculate the linear combination of these two vectors with this specific coefficient, this will be exactly equal to X. And this proves that independent what kind of vector we have in our R2, we can always express that as a linear combination of the vector V1 and V2. And this proves and this concludes our proof that span of V1 and V2 is the entire R2. All right. So we are very close to finishing up this unit. So the next topic we are going to talk about is a linear Independence, and all this important stuff that we learned as part of the previous modules are going to become super handy as part of this specific concept. So we just spoke about the idea of span, we have plotted a lot of vectors, we have seen the linear combination of that and how we can find out whether the span of multiple vectors is the entire space uh for instance the R2, or it is just the line, or it's maybe the zero Vector. We have seen many examples and many operations we have, we have also seen this idea of unit vectors, and we are finally ready to come to this very important concept, which is a concept of linear Independence. So by definition, linear Independence says that the set of vectors is linearly independent if no Vector in a set can be written as a linear combination of the others; otherwise they are linearly dependent. So vectors V1, V2 up to VN are linearly independent if and only if the only solution to the equation C1V1 + C2V2 + ... + CN VN = 0 is C1 = C2 = ... = CN = 0. In other words, in a linearly independent set, the equation C1V1 + C2V2 + ... + CN VN = 0 has only the trivial solution where all Cis are zeros. So now what do we mean here? There is a ton of information in this definition. So let's unpack them. Firstly, it's really important to uh keep in mind this idea of Independence and dependence. Independence and dependence are things that we commonly use in data science, in artificial intelligence, in statistics, so those are really important. So we basically have linear independent condition. So there is a certain condition that our vectors should satisfy, vectors in our set in our Vector space for them to be named as linearly independent, and otherwise we are calling them linearly dependent. And you can see that here there are a couple of Parts as part of this definition. First, it talks about um being unable to create a vector in the vector set while using the remaining vectors in our set. So it says if you can use the remaining vectors in your vector space and linear create a linear combination of them, so linearly combine them, and we have already seen the definition of linear combination. So if we cannot create such linear combination from the remaining vectors to get our Target vector, then we are saying that we have a linearly independent vectors. So if we want to say that all our vectors in our Vector set, they are linearly independent, it means that each of those vectors we should not be able to recreate out of the remaining vectors. So we should not be able to find coefficients to create linear combination using the remaining vectors in order to get our Target vector. Now what do I mean by this Target Vector? What do I mean by this linear combination? Uh I will come to this in a bit. For now let's just try to unpack this definition cuz uh with examples uh we will definitely go through this step by step in detail such that this ideal linear Independence and dependence is super clear. So in the second part of the definition, it says vectors V1, V2 up to VN are linearly independent if and only if the only solution to the equation, and we have here in the left hand side you might recognize the linear combination of our vectors V1 up to VN. So in in the right hand side you have zero. So you are saying our linear combination of vectors is equal to zero if and only if C1, C2 up to CN is equal to zero. So linear Independence basically claims that we will have linearly independent vectors only if and only in the condition when um the only way we can create linear combination of these vectors equal to zero only if those coefficients are zero. There is no other way that we can get a linear combination that is equal to zero while those coefficients are not zero. So the only way that we can get a linear combination out of all our vectors equal zero is only when all of the coefficients C1, C2 up to CN is equal to zero. That's something that we will come later to this again. This something also that we are going to come back in our next module and the next one. So um this one will be also super clear once we go through those modules, but for now keep in mind that the uh linear combination of all these vectors can only be zero in case when all these coefficients are equal to zero. So and then we have the third part in our definition, which says that in other words, in a linearly independent set, the equation C1V1 + C2V2 + ... + CN VN = 0 has only the trivial solution where all Cis are zero. So this explanation is basically what we just spoke about as part of this second part where we said that only in case the coefficients are all zero we can have a linear combination of our vector V1, V2 up to VN which is equal to zero. And why we would like this linear combination to be equal to zero because it's a common way to find solution to our linear system. So this is something that we will also see as part of the next module when we'll be discussing the idea of solving linear systems. We will go into more uh Advanced topics, but for now in order to understand this idea of linear Independence, we should just keep in mind that we cannot find any Cis, so C1, C2, so any coefficients that is not equal to zero and then expect that the linear combination of these uh linearly independent vectors is equal to zero. So that's the uh if and only uh if and only uh if part, which means that this holds from both sides. On one hand we have V1, V2 up to VN which are linearly independent only if the linear equation, so the linear combination of all these vectors is equal to zero if all these coefficients are zero, but also the other way around holds as well. So if we have a linear combination that is equal to zero only if those coefficients are zero, that means that we are dealing with a linearly independent vectors. This is the if and only if part, which means that we have this uh conditions from both sides. If one holds, the other one holds, but also the other way around. All right. So let's now look into specific examples that will make our journey in understanding linear dependence much more convenient. So let's say we have our coordinate system and we have these two different vectors. So we have Vector, let's say (2, 3), which is our Vector a, and we have a vector b that is equal (6, 9). So those two are our vectors, and what we want to understand is where those two vectors are linearly independent or linearly dependent. So one thing that you can quickly notice is that b looks quite similar to a in terms of its scal. So there is a way that we can recreate Vector b by using Vector a. Now you can see that if I take Vector a, which is equal to (2, 3), if I take Vector a and I multiply it by three, so 3 * Vector a, this is a scale multiplication, then what I can get is 3 times and then I have here (2, 3), and this is then equal to (3 * 2 is 6, 3 * 3 is 9). This gives me (6, 9), which is our uh Vector. Now another thing that you can notice that that is exactly my b. So you can see that those two are similar, which means that 3 * a = b. Now what this means is that I can recreate Vector b by using Vector a. So in our definition we saw that a set of vectors is linearly independent if no Vector in the set can be written as a linear combination of the others. So here I can take this 3a as a way to write down a linear combination, so 3a + 0 * b is then equal to b, which is basically saying 3a = b. So by using these two vectors in a set, I can then create a linear combination of the two, and actually even basic way of writing this is saying I can use the vector a to write a linear combination from this, so 3a is a linear combination, so just a scaled multiplication in this case, of course, but if we have just two vectors, our Target Vector is b, and I want to write this uh I want to see whether I can rewrite the vector b as a linear combination of the remaining vectors, which is Vector a, so I can then write Vector b as a linear combination of vector a because I can say that 3 * a = Vector b. So this means that Vector a and Vector b, they are linearly dependent. This means that I can use Vector a to recreate Vector b, and of course I can also do the other way around, right? What I can do is that I can just take Vector b, so I can take Vector b, I can multiply it by 1/3, 1/3 is real number, so I'm just performing a linear combination using b, and this will give me (6 / 3 is 2, 9 / 3 is 3), and I'm getting exactly what I have under a. So I can then also rewrite Vector a by using Vector b. So I created a linear combination using Vector b in order to get a vector a, and that's exactly the opposite what we have learned here because we should not be able to write this a vectors using the other ones in our set, cuz otherwise we have a linearly dependent set. Therefore, we are saying that Vector a and Vector b, they are not a set that is linearly independent, but they are linearly dependent. Before moving on to another example, I also wanted to visualize these vectors just to see what is going on with this pan and uh how the two linearly dependent vectors look like in R2. So this is our R2. We have a vector a which has (2, 3) elements. So we know already the magnitude and the direction. This is 2, this is 3, which means here let me actually use another color, so (2, 3), which means this is my Vector a, and then my Vector b is simply (6, 9). So it is this one. So you can already see what is going on. So this is Vector a, and this entire thing is Vector b, and you can see that those two vectors, no matter how I combine them, I can I will always get the combination, so linear combination of the two on this line. If I want to get um Vector that is for instance in here, I can never find a scalers of C1 and C2 in such way that these vectors, so a and b, they can form a linear combination that will give me this Vector. There is no way that I can do that, and that's why uh we say that this span of this two vectors, so span of a and b with this a and b is this line, and we cannot express any of these other vectors like this one or this one using a linear combination of these vectors a and b. The only linear combinations that we can recreate using these vectors a and b are on this line. So you can see that even if I have two different vectors, I actually just got um single Vector because I have (2, 3), and both of these vectors they are actually um uh scaled multiplication of the other one. So b is equal to I'm missing here something 1/3, so b is simply equal to 3 * a, and then a is equal to 1/3 * b. So in both cases they are simply a version of scaled multiplication of this Vector (2, 3). So a is simply equal to b * 1/3, and then b is equal to 3 * a, and both of them they are actually based on this Vector (2, 3) on this Vector a. So therefore they both actually form and they span or round this single line, and they are both linear. We also call it collinear, and they are linearly dependent. Okay. So let's now move on to the next uh example where we will have bit more interesting case and we will look into this example when we have linear Independence. Look into another example, bit more interesting one, as we want to see whether those two are linearly independent or not. So the first Vector that we got is the vector a, the vector a is equal to (6, 0). So it is this Vector. This is Vector a. The vector b, it is this one, and it contains element of (0, 7). So it is this Vector. This is the vector b. Now in our definition of linearly independent vectors, we saw that the idea of linear Independence is that the two vectors can only be linear independent if we cannot rewrite one of them by using the other. So this means that we cannot rewrite a in terms of b, and we cannot rewrite b in terms of a. So there is no way that we can scale the vector a to get Vector b, and there is no way that we can scale Vector b with vect with some uh scaler in order to get the vector a. So there is no way that we can create a linear combination of this one vector to get the other one and the other way around. So let's see whether this is the case just from uh trial and error. We have a vector a which contains elements (6, 0). For us to go from a to b that has elements (0, 7), so we need to go from 6 to 0 in this case, and we need to go from 0 to 7. Now we can automatically already see from the second element that there is no way that we can go from 0 to 7. You cannot find any scaler c that you can multiply with zero in order to get seven. There is no way that you can do that because any number, any real number that is a real number, if you multiply it with zero, it will never become seven. And of course another thing that you can notice here also very quickly is the other way around, right? So here if you go from this 0 to 6, there is no way you can go from this 0 to 6 because there is no such c that you can take this 0 and multiplying it with that. So here our scaler, and you get this equal to 6. This is just not possible. So what we are seeing here is that there is no way that we can somehow change this vector. So there is no way that we can scale them in such way. So this is -b, so all the scales scaled version of this or all the um scaled multiplications of vector b they will always be on this line, and then the same holds for a as well. So all the scaled multiplications of a will be on this line. So then one thing we can quickly see here is that given that those two are perpendicular, this pen of a and b is the entire R2. So we can see that by using those two lines we can recreate any other line in this R2, and this is highly related to this idea of linear Independence. And given that we cannot come up with a linear combination using the a vectors to recreate the other one. In this case, given that we cannot recreate a using b, and we cannot recreate b using a, so no linear combination that exist that we can use to recreate b using a and the other way around, we are saying that Vector a and Vector b are are linearly independent. Let's now look into another example that will uh clarify this linear Independence concept. So we have three different vectors, and the first Vector is Vector a (1, 0, 0). Vector b (0, 1, 0), our second vector, and the third Vector (0, 0, 1). You can notice that we are in R3.

And then the example goes on and it says that those three vectors are linearly independent. And as an explanation, we have that there is no way to add these vectors together with any scalar multiples to equal the zero vector unless all scalers are zero.

Now, before even going on to the next part, it's actually very quickly, um, uh, provable that those three vectors are linearly independent, and you cannot create a linear combination of one using the remaining two. Let's look into this example in more detail. So we have three vectors, A1, A2, sorry B; so we got A, B, and C, which are (1, 0, 0), (0, 1, 0), and (0, 0, 1).

Now you can quickly see that if we are in R3, and this is actually our unit vector E1, this is our unit vector E2, and this is our unit vector E3, because in that positions we got our ones and the remaining they are all zero. And this is actually very similar to the previous example because we can quickly see how we are, we will not be able to recreate one vector using the other ones by even looking at the positions of the zeros.

So for us to recreate vector A, which is equal to (1, 0, 0), it means that we should be able to find a linear combination C1, C2, and then using these vectors; this is the vector B (0, 1, 0) + C2 * vector C, which is (0, 0, 1). So in here, basically, we are already seeing a problem because we have here, here an element one, and we somehow need to be able to find C1 and C2 in such a way that 1 is equal to C1 * 0 + C2 * 0. But we know that there is no C1 and C2 that we can find such that this uh expression actually is true, because C1 and C2 they should be real numbers, and there are no real numbers that we can find to multiply with zero such that this will end up to one, because this is always equal to zero, and we basically get 1 = 0, which is not true.

And of course, the same holds the other way around. You can prove that B can never be um recreated by using the linear combination of A and C, and also the C can never be recreated by using a linear combination of A and B. Therefore, we are saying, given that A can't be written as a linear combination of B and C, B can be written—so the same, only this time A and C—and then C can't be written as a linear combination of A and B; those vectors A, B, and C, they are linearly independent. And if stronger, you can actually go ahead and prove that the span of these three vectors is the R3, but that's outside of the scope of this example, so we will just pass, but I will leave that um to you to prove. All right.

So now when we are done with that, let's actually move on to the last module, which is the dot product and its applications. So uh the length of a vector and dot product is a concept that um we um are familiar from high school. So the length of a vector is deeply related to this dot product idea. The dot product of a vector v with itself gives the square of the length of V. What basically um it means is that this dot product of vector v—so this thing, which means take the vector v and multiplying it with the with the other vector v—is simply equal to the square of the length of B. So this is a way to express the length of the uh of the vector B, and once we square that, that is the dot product. So that's basically this definition what is about. So we know what this definition of the distance is, and we define it by this, and then we take the square of that distance, and there is our dot product, and we are going to see this idea of dot product a lot, especially when it comes to uh matrix multiplication, vector multiplications, also in many applications of linear algebra you will see this idea of that product coming again uh and coming back to us. So uh this is a concept that we really need to understand.

So in the two-dimensional space, let's say we have a vector B which is um consisting of the two elements X and Y, then the dot product and the link are related by V ⋅ V—this is the way we denote the dot product—so we just simply use the dot, and the name also makes sense because we are saying we are using the dot to perform dot product, so we are multiplying two, we are creating the product of this vector with itself, and this is equal to X² + Y², which is equal to the uh um squared of the distance of this vector.

Now you might recall from high school that we have learned this idea of distance. So if we have x-axis, y-axis, then we basically use this uh X² + Y² to uh get the uh, you know, the formula for from our circle, and then uh we have the X² + Y², we take the square root of it, and then this is our distance. So once we take the square root of that, square of that from this uh square root of X² + Y², then we are simply getting this to cancel out, which is equal to X² + Y². So this is highly related to this idea because we are again talking about distances, and we are simply taking the distance, we are squaring them up, and then we are getting the dot product. So this double uh straight line—this is just a notation that we use—and we spoke about this also before; this comes um from pre-algebra, and this um this is highly important, related to this idea of Pythagorean theorem and how we compute the distances.

So for instance, when we have this uh square triangular—so we have this um uh rectangle here, and we have here the 90 uh uh great—so here we have the right uh right um angle, and here we have our C, which is uh the side right in front of this uh 90° angle, and here we have the A and the B, and we say that the C² is equal to A² + B², and if I were to actually write this in terms of X and Y, so if this side is X and this side is Y, and this is my Z, let's say, then Z² would be equal to X² + Y², and this is something that we can see here too, and the two terms are highly related. So the Z² is equal to X² + Y², and this is simply equal to Z * Z, right, and this is something that we know from high school.

Welcome to module one of this new unit, when we are going to talk about about matrices as well as linear systems. So those are all fundamental concepts that you will see time and time again when applying linear algebra, not only in mathematics texts but also in applied sciences like data science, artificial intelligence, when training different machine learning models and trying to see what is this mathematics behind machine learning models, different optimization techniques when you want to solve different problems using linear algebra.

So in this first module, as part of foundations of linear systems and matrices, we're going to introduce this concept of linear systems, and then we are going to talk about the general linear systems. We are going to uh see this common labeling of the coefficients, this idea of indices that refer to the rows and the columns, we are going to see what is this differentiation between homogeneous and nonhomogeneous systems. So without further ado, let's get started.

So uh the linear systems form the uh bedrock of linear algebra, modeling this array of problems. Thanks to this advancements in these linear systems and solving it in computing, we can now solve a large amount of problems in a very efficient and a fast way. So uh the general linear systems can be represented by this uh set of M equations with N unknowns. In the previous unit, when we were looking into this uh linear combination of vectors, we saw this notation which was A1, and then we we had C1 multiplied, or rather let me keep me uh let me keep the same notation, so we had this linear combination of vectors, so we had β1 and then we had A1 + β2 and then A2, and those are all vectors + A3, so β3 * A3 … and then βm * Am. This is the notation that we saw before, and we said we want to come up, we wanted to come up with the linear combination of these different vectors A1, A2, A3 up to A, and then we use that in order to get a sense of whether we are dealing with linearly independent variables, vectors, or linearly dependent vectors, and then we also commented on the span that these vectors take.

Now, when it comes to um the uh vectors and just in general linear systems, we can represent what we had before now in terms of with a bigger system, so in terms of M equations and with N unknowns. So here what you can see here is that we have M different equations, so we have β, B1, B2 up to BM, so you can see it in here, and then each of these equations it contains N unknowns, so you can see that the unknowns stays the same, so the unknowns are those X1, X2 up to XN, so X1, X2 up to XN are the set of all N unknowns, and then M equations that you can see in here are all these equations: a11X1 + a12X2 … and then a1N and then XN = B1. And here one thing that is really important to keep in mind is that the indexing is what we need to focus on, so we need to keep this one in mind, this aij and this Xi. So this is something that we also spoke about when uh discussing the linear combination of vectors, we slightly uh touched upon on this topic. So let's now dive into this this indexing and how do we index aij? What are this A's? What are this J's? And here you can see that we have a11 and then a12 and then up to the a1N, and this is in our equation one, and then we have in our equation two, a12—let me actually write this with different color—so in our equation two we got a21, a23 up to a2N, and this A's that you see here, those are just real numbers. So a11 can be 1, a12 can be 3, a1N can be 100, and then the same also holds for this B1, for this B2, and for this BM, and all these values A's and B's, they are just real numbers. The only unknowns that we got here are those, so the X1, X2 up to XN. All right.

So what about the indexing now? So we got aij, and as you can see in this case, the first thing that we can see here it stays everywhere the same, which is the one, so we got here one, we got here one, and up to the point we got here one, whereas the second index, this one, it does change, it grows gradually with one, and it becomes, it goes from 1 to 2 and up to N. So you can see here that the first index, first index or index I, it goes from one, it doesn't change, it's just one, so it is 1, 1, and 1. So here in all cases for this equation, I = 1, but another thing that you can notice here is that the index 2, unlike index I, so the second index which is the J, so you see here that the second index is referred as J, this is a general way of defining the indexes, so here J = 1, 2 … and then N. So basically the I doesn't change in the same row, but the J changes, and then of course we have slightly different in terms of I, but then the same for J for our second equation. So here I = 2, and then J is again equal to one and then two … and then N, and then here up to for the last equation, our I = M, and then our J is again equal to one till two … So you might notice that I was looking at this from the row perspective, so I was saying per equation or per row, the I doesn't change, but then the J stays the same, and then it is either one, two up to N, but the set is the same, so it is, it contains all these different elements here, so 1, 1, 2 and then N, but it contains all these different real numbers going from one till N, because we are combining and we are creating this combination, the sum of all these values a11 and then X1, a12X2, a1NXN, and another thing that you can also notice here is that here with the second index, so with this J, J = 1, then here the X's corresponding index is also one. When the J = 2, then the X's corresponding index is also two, and then here the same story, and you will notice that while the coefficient contains two indices, 1, 1, 1, 2 or 1, N, which are the two indices for the coefficients for the unknowns we got, but just single index which goes from one till N. So basically for a, for the coefficients, so I—let me write with the right color—so I can be 1, 2 all the way to M, whereas in case of J, it can be 1 to all the way to N, and the indices are basically used to help us to keep track of in which row we are and what is the um variable that the coefficient belongs to, because knowing this second index, this helps us to understand that we are dealing with a coefficient that corresponds to this first unknown, the first variable X1, and then the same holds in here as you can see in here and in here we are dealing with the same variable X1, therefore the second index, the index J is then the same both in the first equation and in the second one; in both cases it's equal to one.

Okay, so now when we are clear on that, let's also understand this high-level concept, because you will see this system of linear systems, this M equations and N unknowns appearing a lot, not only in terms of calculating and finding the solution to this linear system, but this actually has a very common application when it comes to um running regression, linear regression specifically. And one thing that you can notice here is that here we got also this B1, B2 up to BM, and you will notice that here the index also uh goes from one, but then this time to M. So when it comes to the rows, we have M rows or M equations, therefore we also expect when it comes to counting from the top that at the bottom we will see an M, whereas if we count from this side, so kind of like imagine it like a column, then we see that it goes from one till N. So those are common observations and reference to um number of observations and number of uh features that you will see in your data when dealing with data analysis or modeling data. So just this uh just keep those things in mind, this uh abbreviation of M and then N, M equations and unknowns, because this will become very handy. And the same also holds for this indexing, just to keep in mind that this I and this J, what those indices are, and how, for instance, the first, you know, the I, the first index changes when we go from up to the bottom, and how the second index J goes and changes when we go from left to the right when we go through the columns, but we are going to, it is also in the uh upcoming slides, so uh we can we will have time to practice it. So um this is what we are calling a coefficient labeling, the coefficient aij. So this thing in a linear system, they are labeled where the first index represents the row and the second index denotes the column. So when we see aij, we know that this refers to the row and the J refers to the column. So this is something that we use in order to understand where exactly in our matrix—something that we can we will see very soon—where exactly our unit or our uh member that is part of our matrix, where exactly is that located, in which row and in which column. The systematic labeling is super important because this helps us to keep the structure and this helps us to understand uh what does this uh coefficient represent, what what is this row that it belongs and what is the column it belongs, so for which equation and for which unknown we have already solved the problem such that we can know what this uh coefficient represents.

So before moving on onto the actual linear systems and the definition of matrices, let's quickly understand this distinction between homogeneous and non-homogeneous, because this will help us to also get an understanding how we can solve a system of linear systems. So a system is homogeneous if all the constant terms bi are zero; otherwise it's non-homogeneous. So identifying this helps us to really understand the nature of the solution set that we need to get and to understand what kind of strategy we need to use in order to solve this problem. Now what do I mean by bi? We is so that we had this system of M equations with N unknowns, and we saw that that we have in the right hand side this B1, B2 up to BM, which means that we had this M different equations with N different unknowns, and to find a solution to the system it means finding this value, values corresponding to X1, X1 here, X2, X2, XN, so basically finding the set of X1, X2 up to XN that solves this problem, and for us to know how to solve this problem, we need to know whether this B1 is equal to zero or not, this B2 is equal to zero or not, and then this BM is equal to zero or not. This is very similar to this idea of solving any sorts of um problems that contain unknowns; for instance, if we have 3x = let's say 5, solving this is entirely different than if we know that the 3x = 0. So this is a simplified version of course, but the idea is the same. Knowing that this B1, B2 up to BM, this R zero, this gives us an idea how we can solve this problem, and later on we will see this distinction between non-homogeneous and homogeneous system, and whenever these Bs, so whenever this B1, B2 up to BM, whenever these Bs are zero, then we are saying that the system is homogeneous and we need to solve a homogeneous system; otherwise we are dealing with a nonhomogeneous system, so this means that the bis are not all zero.

Let's now move on to the second module, which is about the matrices. So we are going to define the matrix, we are going to see the definition of it as well as the notation, this idea of rows, columns, dimensions, uh some of which we have already touched upon, but we are going to uh go into the depth of it, we are going to learn properly as well as we are going to see many examples. Then we are going to talk about matrix types. So here we will talk about identity matrix, diagonal matrices, and also special type of matrices like matrices containing only zeros and only ones. So by definition, a matrix is a rectangular array of real numbers that are arranged in rows and in columns. For example, an M by N matrix A can be represented as follows. So let's look into this definition and this reference to matrix. We call this matrix or matrix A, and every matrix it can be described by this rows and columns, where we always have this uh way of describing this matrix, always should be defined by the number of rows and number of columns. So this is super important, and let's look into this specific matrix. So we have a matrix A, and all these values, they are members of this matrix, they form the matrix, and we already saw this labeling of aij, where we said that I is referred to the row, so you might recall that those were all these equations that we got, so this horizontal lines where I = 1, I = 2, I = 3 up to the point of I = M, and then we had this J, so this thing, and then J was referred to the columns, and we

Had J was here one and then two and then three up to the point of n. So one, 2, 3, and N. This is exactly what you can see here. So in this Matrix, we got all these elements: a11 is a number, a12 is a number, up to the a1n is a number. Those are all real numbers.

And one thing that you can notice here is that here we got a11, so this is our first row and first column. Here we got a12; this is our first row and second column. And then we got up to the point of a1n. Actually, let me just write this down even at a bigger scale such that I can make more notes. So let's assume we have this Matrix a, and this Matrix a is bigger, and we got all these different elements. So we start with our first row, and here we have A11. So here the row that I will write with, let's say, with blue, the r is equal to 1, and then the column is one. So this is Row one, this is Row one, and this is column one. Let me write it with red: this is column one, this is column two, this is column three..., and this is column n. And this is row two, this is Row three..., and this is row M. So in total, I got M rows and N columns. I will come to this notation that I'm putting here later; for now, let's keep track of the rows and the columns to get a good understanding what this indices were about that we just learned. So every time I will also mention this reference to aij to keep track of this. And also, let me write it with the right colors. So ai, this is the row, and j, which is the column. So all the elements I'm just defining by this a because it's just a way to reference a part that comes from a matrix; it's a just common way to write the higher matrix by capital letter A, whereas its members we will write with the um, with the lower case a. So this is Matrix, Matrix a, all right.

So here in the second row but first column, we got a21 because it is still in the First Column. And then when it comes to this element, we have here a; the row is the first one because we are in the first row, but then we are in the second column, so this one should be two. Then we go on to the next element in our first row, so a1, and then three, and then... the last element is an a, as we are still in the first row, it will be one, the i, but then, given we are in the last column, the column index or the j will be equal to n because we got in total n columns. So we are now ready to go into the second row. So here, given that we already have our first element a21, this is in our second row and the First Column, so the i is equal to here two, and j is equal to 1. Let's now write down the element in the second row, second column, as you might have already guessed: i is equal to here 1, i is equal to here 2s, and then uh, the j is equal to 2. And then we go on to the next element, which is in the second row and the third column, so it's a; the row index is 2, so i is equal to 2, and then the column index is three..., and then we got a, as we are in the second row, it is the i is equal to two, and as we are in the last column, the j is equal to n.

Now you might have already guessed when I was writing this down that whenever you are in the row and you move on to all the elements in the same row, the i, so the row index, it stays the same; only you need to uh update the column index. So here, for instance, you got one, one, one here also one, so all the way down in the same row are one, which logically makes sense because we are in the same row, so the row index should not change, but instead you should change the column index, like here: column one, column two, column three, all the way to column n. So those are our columns..., so let me make this distinction, and those are our rows, as you can see. So this kind of mentally helps us to understand why we are writing all these indices over time. Once you practice more with this, this will become more natural very quickly. Remove this, so now our ride rest very quickly. So as you might have already guessed, we are in the third row, so we have a tree. So everywhere I will just write down the a's, so first write down the A's, and then the rows; the row index will stay the same as i in the same row, but then I will increase the columns gradually. So we are in the column one and the column two, column three, up to the column n. So now the remaining stuff you can actually write down yourself to just practice. Let's now move on to the last row and last column. So in the last row, we got a, a, a up to here, and in the last row, the uh row index is M, which means that here I need to have M, M, M everywhere; I need to have M, and then the column index is 1, 2, 3 all the way to n.

So this last column is very interesting too. You can see here that we have the opposite of what we have here because in the last column we see that the uh column index is the same, so it is everywhere n. Only the first index, the index of the row, it changes; it goes from 1, 2, 3 up to M, which is of course logical because we said that in the last column, if we are looking it from the perspective of column, so all these values, this A's, so the all the ends, they are logical because they we are in the last column; we are in the same column, but then the row changes. Here we are in the row one, here we are in the row two, Row three up to row M; therefore, we have also at the end amn.

Now let's talk about this idea of MN. We said that our Matrix a has M as a number of rows and n as a number of columns, which you can see by the way also here. So we always refer the dimension of a matrix, so the dimension, dimension of Matrix a by these two numbers. So first we always write down the number of rows, in this case M, then as the second element we are writing the number of columns, in this case n. We are always putting this small x in between to kind of emphasize M by n Matrix, and we most of the time use the square braces to showcase that we are dealing with dimension, and in this case we are saying the dimension of Matrix a is equal to M by n. So we are dealing with M by n Matrix. This is a common convention used in linear algebra in mathematics general, but also used in data science, uh in machine learning, artificial intelligence. So whenever you are dealing with matrices a, it is a common convention to talk about this idea of dimensions, and the idea of Dimensions is super important when it comes to the idea of multiplication, multiplying Vector with Matrix, Matrix with Matrix, so this dot product. Dimensions play a central role in here, so keep this one in mind. Once we uh get to the point of dot products, this one will become very handy.

So let's now look into a specific example where we see simple Matrix a. So in this case, you can see that we are dealing with a matrix that has a 2x3 Dimensions. So like we just learned, 2x3 means that we got two rows and three columns; that's something that you can also see here very quickly. So you have a small Matrix; on the small matrix it's really easy to actually count. So you can see that we got Row one and row two, and we got column 1, column two, and column three. So this basically confirms this Dimensions; therefore, we are also saying that we have a 2 by three Matrix, and like usual, we first write down the number of rows and then the number of columns. You can see here that here we have this elements for our Matrix, so a is equal to 1, 2, 3 for the first row, and then uh 4, 5, 6 for the second row. So from this, actually I think it's a good exercise to just uh vary our understanding of indices, and from this um we can write down that for instance all these different elements uh like a11 is equal to 1, a12 which means that we are in the first row and in the second column, so we have this element is equal to two, and then we got a and then one three, so we are in the third column, so this one 1 is equal to 3, and then a21 is equal to 4, a22 is equal to 5, and then a23 is equal to 6. So this is actually a good way to practice our understanding of indices, our understanding of this Matrix structure, and the understanding of dimension of the Matrix, which in this case is 2x3. So this is yet another different definition of a matrix structure when it comes to the rows, columns, and dimensions. So this is exactly what we just spoke about on our example, and let's just quickly look at the formal definition. So the rows of a matrix are the horizontal lines of the of the entries, while the columns are the vertical lines. So basically it's saying those are; let me remove this. So the rows are are the horizontal line, and the columns are those vertical lines; those are the columns. This helps us to form these columns, so column one, column two, and column three, whereas this horizontal lines it helps us to create the rows, so Row one and row two. So then we have the dimensions of Matrix are given by the number of rows and columns it has. So an M by n Matrix has M rows and N columns; that's something that that we already saw.

Let's look into some special type of matrices. One Matrix type is the identity Matrix. So we saw before we had this Identity or unit Vector; now we have identity Matrix. So the two are quite similar. So like before when we had our unit vectors, we had this for instance E1 in three dimension; we had 1, 0, 0, then we had our E2 which had 0, 1, 0, and then we had our E3 which was 0, 0, 1. So you might recall this about our identity vectors, or we were calling it unit factors. You might notice very quickly that we have formed an identity Matrix in, which is a square Matrix with one on the diagonal and zeros elsewhere. Is basically a matrix that is built using those unit vectors. So here we have E1, here we have E2, and here we have E3. So you can also see that this 3x3 Matrix because we got three rows and three columns. So you can see that here we have on the diagonal; so we call this diagonal; on this diagonal we have all ones, and in here, outside of the diagonal, they are all zeros. And this is the definition of identity Matrix; it is this in Matrix where n is the dimension of a matrix, and given that it's a square Matrix, it means that the dimension of it is n by n. So all the rows, so the number of rows is equal to the number of columns; on the diagonal we have all these ones, and everywhere else we got zeros. And do note that we are forming this identity Matrix simply by combining these different uh unit vectors, so like here E1, E2, and E3. So let me actually uh give you yet another example, but of much higher Dimension, so of this identity Matrix. So let's say we have I, and then this I uh let us actually use this notation in. So let's say we got in; what this means is that we got actually this large matrix; it's a square Matrix, which means that it is n by n, so it has n as the number of rows and n as number of columns. So the dimension is n by n; you got n as number of columns too because it's a square. And let us actually write down that how that Matrix looks like; it's a large Matrix; the N is the size of that Matrix. So here we got on the diagonal we got one here, we got one here, we got one..., up to the last point one, and the index of this one here, so this is the first row, this the First Column basically, and everything else is simply zero. So here we got z0, 0..., zero; here we got 0, 0 all the way down to zero; here also zero all the way down to z, and then here also zero. So everywhere here and here we all got zeros; only on this diagonal we actually got ones. So basically by using our common notation, we can say that in the d in the identity Matrix we got a11 = to a22 = to a33 equal to all the way to ann equal to 1, and then when it comes down to the rest of the cases, so all the other observations, let's say a21, a31 or a41, anything, so anything that is not um a11 or a22, anything that is not on the diagonal, it is simply equal to zero. We also say in those cases that aij is equal to 1 if i is equal to j, because then it means that we are talking about item that is on diagonal because both the row index is equal to the column index; otherwise, the aij is equal to zero if i is not equal to j. So this is in the nutshell how a large identity Matrix in general can be defined.

So let's now move on to another type of Matrix, which is the diagonal matrix. So by definition, a diagonal matrix is a matrix where all off-diagonal elements are zero. So what does this mean? We saw um example of a diagonal matrix, which was our identity Matrix, because identity Matrix is an example of a diagonal matrix. And what do I mean by that? In our just seen example, we saw that only on the diagonal we had all these nonzero elements, but the rest were all zeros. So all the off-diagonal elements were zeros, like in here and in here. Exactly the same holds for the diagonal matrices; only, unlike in the identity Matrix, we no longer need to have this diagonal elements equal to one; those can be any other numbers. So as long as we have this um elements D1, D2, D3 that are not zeros, but then of the diagonal numbers, so all these elements, they are zero, then we are dealing with the diagonal matrix. So in this case, we got a 3x3 diagonal matrix because we have uh three rows and three columns, and here we can see that the um the first, so the a11, the first element from the first row and First Column is equal to D1, so a22 is equal to D2, and then a33 is equal to D3. So D1, D2, and D3, those are all; so D1, D2, and D3, those are all real numbers. Now when it comes to the uh this numbers, for example, it can be that D is let's say 2, 5, 6 on diagonal, then we have those zeros; this is a diagonal matrix. It can also be that D is equal to minus 3, and then 0, 0, and then 5, 8, and then here we have zeros. So again, we have on the diagonal all these elements, and the off-diagonal elements. So if all the off-diagonal elements are zero, then we are dealing with diagonal matrix. And if you're wondering, well, what happens if on the diagonal we got zero? Do we still have a diagonal matrix? It's actually a great question, but yes, indeed we are dealing with the diagonal matrix as long as all the off-diagonal elements are zero. So for instance, if we got D is equal; here we have zero, here we have 0, 0, 0, and then 7, and then 0, 0, and then 8, and then 0, 0. So we got this off-diagonal elements, so here are the diagonal elements, and all the off-diagonal elements are those. Given that all the off-diagonal elements are zero, which is the definition of the diagonal matrix, then we can say that our D Matrix in here is indeed a diagonal matrix.

Let's now look into yet another type of Matrix, which is a special type of Matrix, and it's called one's Matrix. So by definition, one's Matrix is denoted by 1M1. So you can see here and here; it means the dimension of it, so the number of rows and number of columns, and it's a matrix in which all the elements are one. So this is a very unique Matrix; we often use it during the programming, so in data science, data analytics, but also in um uh when creating like data structures, when designing algorithms, this becomes very very handy. And this idea of one's Matrix is that all the elements are just one; it means that if we want to create a placeholder in such way that we can then multiply any number in here with some other number and get that number, then it can be done very easily because we know that when we multiply a number with one, then we get that number, so a * 1 is = to a, x * 1 is = to x. Now this is a exactly this property exactly is what motivates us to create and to have this type of ones matrices; it means that we are defining matrix by its Dimension, so it is M by n, and here the m is equal to two, and then n is equal to three because we got two rows and three columns, but you can see that all the elements are the same and they are equal to one. So a11 is equal to a12 is equal to a13 is equal to a uh 21 and is equal to a22 and a23, and they are all equal to one, and this is the definition of one's Matrix. You can have um one Matrix of the size 4 by 10, one's Matrix of the size th let say 10,000 by 100, etc. So any number, any real number, so M and then n are real numbers you can use in order to create this large M by n one's matrix.

Let's now look into our final special type of Matrix before moving on onto the next module, which is about zero matrices. So similar to this one Matrix, a zero Matrix, denoted by 0m by N, is a matrix in which all the elements are the same, with the one difference that this time all the elements are equal to zero. So in the one Matrix all the elements were ones, but in the zero Matrix all the elements are zero. This type of matrices become very handy also during the programming, creating um different algorithms during design, encoding um, but for slightly different purposes. Usually we create the zero matrices as a placeholder such that in the beginning we can have this uh tuples or we can have this um uh arrays or nested loops um that we want to perform, and then gradually add these values to the existing m array. So if we create this zero Matrix and um this is a placeholder, then next time we can always add on this this new data that we get, and then we know that zero plus a number is always equal to number, which means that once we have this updated information of a, we can add this to the zero, and we will then have this new updated information in our system. Therefore, the zero Matrix is often used as a way to uh have this placeholder with the provided Dimension where we can always add new information, and the information can be updated. So in this specific case we got um a zero Matrix that has two rows and three columns, so you can see two rows and three columns, so m is equal to 2, and then n is equal to three. Three, perfect. So we are done with module 2, and now we are ready to go on to our next module, which is the core Matrix operations. So when it comes to matrices, we often perform Matrix additions, Matrix subtraction, but also Matrix um scalar multiplication of this Matrix, so multiplying Matrix with a scalar, and then Matrix um multiplication just in general, so taking two matrices and multiplying them. We are going to look into this concept in detail; we are going to see many examples, like before; we

Are going to dive deeper into this such that we lay the ground on uh, to the next module, which is solving a system of M equations with an unknown. So, solving this general um, linear system. So for the beginning, uh, we will be looking into this matrix operations where we are adding or subtracting matrices.

By definition, the sum of two matrices A and B of the same dimensions is obtained by adding their corresponding elements. So, by taking the element i j from both matrices and adding them to each other. So, in this case, you can see that Matrix A and B are here and uh, the uh definition says we just simply need to take the corresponding elements, corresponding elements from the row I and the column J, take them, add them, and this will become an element in our final um Matrix. Because when we are adding two matrices of the same size, the result is yet another Matrix. So we will use the Matrix a to add to Matrix B, and this will give us a matrix A + B. And this i j simply refers to the indices corresponding to the row and the column. We will look into an example in a bit, and this will make much more sense.

The same holds also for the difference. So, by definition, the difference of the two matrices A and B of the same dimensions is obtained by subtracting their corresponding elements. Which means that in order to obtain this Matrix a minus B, this is a new Matrix, we simply need to look for each element. So we are going to index them for a row I and J; we are going to do this pairwise element-wise subtractions. We are going to see what is that element corresponding to the row I and column J in the Matrix a, which we say is a i j; we are going to subtract from this the element in the row I and column J that comes from Matrix B, and this will give us our new Matrix, which is a minus B. So let's now look into an example. In this Matrix Matrix um uh a and Matrix B are used, and Matrix a is of the size 3x3 Matrix 3; Matrix B is of the size 3x3.

In order to obtain a + b, what we are doing is that we are performing element-wise additions. Now let's verify this. So what we are doing here is that we are saying a + b. Let me actually get a larger area here. So let's say we have the two matrices. I want to add the two in such a way that we do everything one by one such that this idea of a + b and addition of the matrices will make sense. So we want to find out a + b. For that, what we are going to do is that we are going to make use of this definition that A + B and then I J is equal to a i j + b i j, which is a fancy way or mathematical way or describing that for each element we need to go and look for the row I and column J and take that element from the um column from that uh Matrix a and from the Matrix B. So this means that for a + b, this is going to be a matrix that will have the same number of rows and the same number of columns as two matrices because both A and B are 3x3, which means also their sum is going to be 3x3. So this going to be 3x3, and here we are going to do so we are going to take for the first row in the first column. So for A + b 1 1, so first row and first column, we need to go to the first row and first column of Matrix a and the first row and first column of Matrix B, and we need to add these two elements. So we need to do 1 + 1, and then we need to go on to the second column, so the first row and the second column, which means that we need to be here in both matrices. So here we have 0 + 2, and then we got 2 + 3, and then we got 0 + 0. So you can see it in here, and then we have 1 + 0, and then we have 3 + 1, 0 + 1, and then 0 + 2, and then 1 + 3, which gives us so 1 + 1 is = to 2, 0 + 2 is = 2, and then 2 + 3 is = to 5, 0 + 0 is equal to 0, 0 + 1 is = to 1, 1 + 0 is = to 1, 0 + 2 is = to 2, and then 3 + 1 is 4, 1 + 3 is = to 4, which means that our A + B is equal to this Matrix that we got in here. So you can see that we are getting exactly what we uh what we have here; only we have done it manually one by one. So the same idea holds exactly when we have a - B, only instead of adding, you will have to do here minuses, so minus minus, so everywhere minus, so 1 - 1, 0 - 2, 2 - 3, etc. So let's look into another addition.

In this case, by definition, it is defined as this element-wise uh of the adding of these two matrices here. The only difference in this definition is that it's saying it's calling this a + b as C. So this new Matrix that we are getting as a result of adding a to B, it's calling C. So basically, it's the same as calling this Matrix as C. You will see also this type of definitions. So in this case, the Matrix C is equal to a + b, which basically means that for each row with index I and with each column with index J, go and look for row I and index J, take the corresponding elements from Matrix a and Matrix B, add them in order to get that corresponding element in our new Matrix C. And you can see that in this example, that's exactly what we are doing. We have a, we have B; we are taking this element and this one, so 1 + 1, we are getting here two, and then 0 + 2, we are getting two here; 2 + 3 is 5, and then 0 + 0 is = 0; 1 + 0 is = to 1, and then 3 + 1 is equal to 4.

So now we already go to the next topic, which is about scalar multiplication of a matrix. So by definition, scalar multiplication of a matrix a by scalar Alpha results in a new Matrix where each entry of a is multiplied by Alpha. The idea of scalar multiplication matrices is actually quite similar to this idea of scaled multiplication in vectors. So uh, we have already seen in the lecture of the vector multiplication that when we were having this scaler C and we had this Vector a, then uh when we are multiplying C, which is a real number, with Vector a, then we simply need to take all the elements of vector a, so A1 A2 all the way down to a n, and we need to multiply them by this same scaler, so see this is what we were doing with vectors, and that's exactly the idea behind matrices. And when uh doing the scalar multiplication of matrices, only instead of multiplying only just one vector with this scaler C, now we need to apply this to all the rows and all the columns. So here we got this one column, and Matrix is simply a combination of multiple vectors, which means that we need to multiply all these elements of all the vectors of all the columns in this Matrix. So let's actually look into a specific example.

So in this case, we have a matrix a, and this Matrix a is this thing, and we have a scaler which is three. So in here our Alpha is equal to three, or you can call it C or anything. So you can see that when we are scaling the Matrix with a scaler, in this case three, with this Matrix, what we are doing is that we are simply taking each of these elements and multiplying it with this scum. So 1 by 3 is 3, 2 x 3 is 6, 3 x 3 is 9, and 4 x 3 is 12. This is the idea behind this entire scal multiplication ofation Matrix. In more general terms, if we for instance have a matrix a, so let's actually look into a high-level general example where we have a DA Matrix M by n, so we got M rows and N columns, and we want to get a scal multiplication of this Matrix and um scaler that we have here as in our definition; it is defined by this alpha; alpha is just a number, you can qu C, you can qu B anything. So in this case, our scaler alpha, alpha is coming from R, so it's a real number. So Alpha time a is then simply equal to to this new Matrix where all of these elements are simply multiplied by this scal. So I will just take over all these values H1 up to a M1, and then A1 2 a22 all the way down to a M2, and then let me also add the last column just for fun here a 2 N and then here a m n. So here this new scaled M multiplies, so so scaled uh Matrix a, so Alpha * a is simply equal to Alpha time all these elements are simply multiplied by the scale; it is as simple as that. So that's the simple idea behind um Matrix as scaling. So when you are doing scalar multiplication of this Matrix, you simp take all the values and you multiply them element by element per row and per column by that single scalar Alpha. Do note that you are multiplying them all without exclusion with exactly the same number, which is that Alpha.

So let's now look into the definition of matrix multiplication. So here we are no longer multiplying a matrix with a scalar, but we are multiplying Matrix with Matrix. So the product of an M by n Matrix a and an N by P Matrix B results in an M by P Matrix C where each entry cig is computed as the dot product of the e Road of a and the J column of B. Now what does this mean? Firstly, let's look and unpack this part of the definition. So we got Matrix a that is M by n, and then we got Matrix B which is n by P. What this means is that in this case, Matrix a has M rows and N columns, and Matrix B has n rows and P cups. So this is then simply the dimension dimension of the two matrices. So then it's saying that by definition, the product of these two matrices, so the product of A and B, the product of the two B is equal to to this Matrix C, and each entry cig J, so c i j is computed as the dot product of the each row of a and the Jade column of B. Now this part might seem a bit difficult, but once we look into the actual example and we illustrate this on our common high-level general expressions of Matrix am and their multiplication, this will make much more sense. For now, before coming to this one, I just wanted to refresh our memory on one thing I said before when discussing also this idea of improving uh this uh different properties of vectors that when we want to multiply a vector with a matrix or Matrix with Matrix or vector with a vector, we need to ensure that from the first element, the number number of columns is equal to the number of rows of the second element. This is also very important for this specific case and just in general for matrix multiplication. So you can notice here that the number of columns here is equal to the number of rows in here, and the order is very important. So in case of matrix multiplication, the order is really important, which means that if you have a matrix a and you want to multiply with the Matrix B, then the number of columns of a should be equal to the number of rows of B; otherwise, you cannot multiply those two matrices with each other. So in case you got a matrix a that doesn't have the same number of columns as the rows of number of the Matrix B, then there are some alternative things that you can do, including this idea of the transpose that we saw also doing when computing the dot product between this Vector a and Vector B. That's something that we also do in programming when we are dealing with this Matrix and we want to compute this relationship between two matrices, but the number of columns of one of the first one is not equal to the number of rows of the second one; we are simply uh manipulating this matrices or removing some data if that's not hurting our problem, maybe uh flipping, so transposing our Matrix or applying any other source of operation to it to ensure that the two matrices that we are multiplying with each other, the first one's number of columns is equal to the second one's number of rows. That's just the low, and that's something that you should follow if you want to multiply these two matrices. All right, so now let's move on onto this idea of multiplying and dot product. Let's look into a specific example, and this will uh help us to understand this process better.

So before doing that, I just want to quickly show you this general idea. So if we have a matrix a that is M by n, which means that it looks something like this, like A1 1 a 2 one up to the point of a M1, and then here we got let's say a 1 2 a 22 up to the point of a M2, and then at the end we got a MN, and here we got A1 n. So let me also add this one 2 N, and we got a matrix B. This Matrix B is n by P, so it has n rows and P columns, so we are fine in terms of Dimension here, and we got here b11 B21 up to the point of b m, sorry, b n in this case, let's not confuse the letters, so b n 1 B1 2 B 22 up to the point of b n 2, because n now is the number of rows for Matrix B unlike for the Matrix a, up to B1 p and here b 2 p and here after the point of B and then n p; this is the last element. In order to perform um multiplication between these two matrices, so to obtain a matrix C which is a equal to a * B, what we need to do is we simply need to take pair case or pair Row for the row I, for instance, we need to take this element, so this row, and we need to multiply it with this, so we need to find the dot product between this row and this column; then we need to move on on to the next one, and then for the second element, we will then take this row, and we will multiply it with this one. So this is then something that we need to do in order to obtain these elements, and you might have already noticed that we got this m by n and n by P, so you might have already guessed what will be the dimension of the C. If we got that the dimension of a is equal to M by n and the dimension of B is equal 2 N by P, then the results Matrix after M multiplying the two, so Matrix c will be will be having a number of rows equal to this and the number of columns equal to this. So this middle part basically disappears, and the number of rows of the first Matrix will be then the number of rows of this result Matrix C, and the number of columns or the second Matrix, so Matrix B, will then be our final number of columns. So we will then have a matrix C that will have a dimension, so Dimension, so dimension of C will then be equal to M by P, so we will have M rows and P columns. So how we are going to compute this? So for c i j, which means row I and column J, let's look into the definition of it; it's saying c i j is computed as a dot product of the each row and the Jade column, so each row from a and Jade column of B, what where is the each Road of a? The each Road of a is somewhere here, so each Road of a it is uh the A and then I one then a and then I 2 and then a and then I three dot dot dot and then a i and then then we got in total n columns n, and we always do the transpose right when computing this um dot product, so we then take the transpose, so we take this row row I and we multiply it, so we do the dot product between this one, this is the a i, and the B J, this is column J, it is somewhere here, so it is B and then we got the first element which is one and then J and then b 2 J B 3j dot dot dot up to B and then in total we got n rows in B, so n and then the J is the column, so it stays the same. So this is then the dot product between e row that comes from Matrix a and the J column that comes from Matrix B. So it's always like that actually, so we always take row by row, so we take this different, so every time we take just a row and we multiply with the corresponding column, and then we get the dot product between this row that comes from the first Matrix and then the column that comes from the second Matrix in that specific order in order to get our DOT product and that specific value. And what is this amount actually? So when we calculate this do product, you can quickly see that we have a i1 multiplied by B 1 J plus a I2 multiplied by b 2 J and then dot dot dot a i n multiplied by b n g, and this new Matrix c will then have all these elements, so C11 C 21 and then c31 dot dot dot and then C the last row as the number of rows of C is m c m see here M, so C and then here it will be one 2 C 22 C3 and then 2 up to the point of cm and then two, and then here the last col will be C1 and then p is the number of columns in C, so C1 p and then C2 p and then c3p dot dot dot and then c m and then P. Okay, so this is what we get; this is our final Matrix C when multiplying Matrix a and Matrix B. So let me clean this up. C is to now if you want to find out what is C11, you can easily fill in this general formula that uh that we just calculated the I is equal to 1 and then J is equal to 1, and this will give you C11 by using this formula. If you want to get the C and P, then just fill in the I is equal to M and then J is equal to P in order to get this value C and P. So you can already see the amount of calculations you need to do in order to get all these elements from this large matrices A and B. Let's actually look into a simple example to clarify this.

So we have a matrix a here and Matrix B here, and we want to do a multiplication of the two, and we have just learned how to do it. Let's actually do it one by one. So we got a matrix a which is equal to 1 2 3 4 with Dimensions 2 by 2; then we got a matrix B which has values two Z and then one two, so it is 2 by 2, and I want to find what is c that is equal to a * B, and I know already by looking at these Dimensions that c is going to be equal to 2 by 2. So you might recall that I said that when looking at this final result, the number of rows or the final um Matrix will be this, so the number of rows of the initial Matrix a, and then the number of columns of this final Vector c will be the number of vectors number of columns of this second Matrix B, so two. Therefore, I know already before even doing calculations that the uh product Matrix c equal to a * B is going to have a dimension 2x2. Let's actually do a calculation to check this. So C is then equal to a * B and it's equal to 1 2 3 4 4 multiplied by 2 0 1 2. Okay, so I expect to have four different elements here, here, here, and here. So to obtain the C11, so it is C11 in here, what I need to do is

That I need to look at the first row and in the first column in here. So, first row from A and the first column of B, and I'm doing the dot product, which means 1 * 2 + 2 * 1. 1 * 2 is 2; 2 * 1 is 1. So here I'm getting 1 * 2 + 2 * 1, which basically gives me 2 + 2, and that's equal to 4. So here I'm just writing [Music] down 1 * 2 + 2 * 1.

Now, when I want to get this value, which is C12, this means that I want to get the first row and the second column, and that's exactly what I'm doing. So I'm going back and I'm saying, let's look at the first row, but this time we'll look at the second column coming from the uh from the Matrix B. So 1 * 0 + 2 * 2. And then I do the same, only this time for the second row, which means I'm picking this row and then this column. So it is 3 * 2 + 4 * 1. And for the final element C22, I'm taking the second row and the second column, which gives me 3 * 0 + 4 * 2.

Now, what does this give me? This gives me this 4x4 Matrix where 1 * 2 + 2 * 1 is 4; 1 * 0 + 2 * 2 is 4; 3 * 2 + 4 is = to 6 + 4, which is 10; and then 3 * 0 + 4 * 2 is = to 8. So let's check: 4, 4, 10, 8. That's exactly what we have here.

So, as you could see here, the idea is that every time to follow what element I'm looking for for the CIJ, and then I just go to the I rows from the first Matrix and the J column from the second Matrix, and I do the dot product of the A and then I and then K, let's say. So I'm going to the I row from the first Matrix, and I'm taking all the elements, which means I don't even need to mention this index; it just means the entire each row coming from the Matrix A. And then I'm doing the dot product between this row and the column that comes from the Matrix B, which means B and then J, which then will give me the CIJ. So I'm looking at this and taking this, multiplying this dot product, and this gives me the first element; then the first row and then the second column, which gives me the uh second element in the first row in my Matrix, so this one, and so on. So hope this makes sense. Uh, if it doesn't make sure to reach out because it's a very important concept, and uh let's also look into another example to make sure that we got this right.

So in this case, as you can see, we have another matrices, so set of A and B matrices again, 2 by 2, a simple one, and we want to know what is A. So let's say we call this C; we already know C should be 2 by 2, and what we are doing is basically for C11, we are saying, let's look at the first row, so first row and the first column coming from the second Matrix B, and let's do the dot product. So 2 * 1, 2 * 1, 4 * 5, 4 * 5. We get this. And then when we want to find what is C, oh, what is C and then 1, 2. So in the first row, but in the second element in our final Matrix, so I is equal to 1 and J is equal to 2, it means we need to look at the first row from the Matrix A, but this time the second column from the Matrix B. So it is 2 by 3, 2 by 3, 4 * 7, 4 * 7, and this gives us a number 13, 4. Even if you calculate, you can see that 2 * 1 is equal to 2, 4 * 5 is 5, so 2, 4 * 5 is 20, so 2 + 20 is 22 in here, and then you can do the rest of calculations, and this will be a good practice to see how we can do a basic matrix multiplication. The idea is actually quite straightforward when it comes to multiplying it; it just it comes with a practice when we see all these uh much bigger matrices. So um this is another example; I will leave this one to you to complete it, just uh to keep in mind we always do uh so we always look at the dimension first in here, 2 * 2 and 2 * 2, which gives me an impression already what I can expect the result will be 2 by 2. And when it comes to the uh cross elements, just ensure to always look to the I row and the J column; this comes from Matrix A, and this comes from Matrix B; take them, compute the dot product, and then you will find your C, your final result; let's call it um CIJ because in this case we have a matrix C already.

Welcome to the module 4 of this course when we are talking about matrices and linear systems. So in this module we are going to dive deeper into this uh idea of linear systems with matrices and solve linear systems using different techniques, and specifically we are going to learn the uh concept behind solving linear systems using matrices named Gaussian elimination and Gaussian reduction.

Welcome to the module 1 in this unit. So in this uh case we are going to talk about algebraic laws for matrices. We are going to discuss four different properties for matrices, uh and the first one is the commutative law for matrix addition, the associative law for matrices, the distributive law for matrices, both the left and the right one, and then finally we're going to talk about the scalar multiplication law for matrices.

So the algebraic laws or matrices, they are like in case of real numbers, like in case of vectors; they help us to do different operations on these entities. They are very similar to the real numbers and the vector cases where we for instance um so that if for instance A + B uh is equal to B + C, or A * um B + C is equal to A * B + A * C. Those are all sorts of laws that we uh learn as part of high school prealgebra, and we have applied it to real numbers; we know how helpful those can be, and similar type of laws we have also for the matrices, and we got in this case four different laws that we will be discussing. The first one is what we are referring as associative law; the second one is the distributive law; the third one is the scalar multiplication law; and the fourth one is the commutative law for addition. So these laws help us to do different matrix operations; they help us to manipulate algebraically these matrices, and then uh this can help us to solve different sorts of problems, including solving a system of linear equations.

So let's start with the commutative law for matrix addition. So the matrix addition is commutative, um which means that A + B is equal to B + A. So unlike the matrix multiplication that we have seen in the previous lessons uh where the order did matter, and we said that we um had to uh ensure that the number of columns of the first Matrix is equal to the number of rows of the second Matrix, in case of addition that's this is not the case. So we should not care about the order whenever we want to add two matrices. The other thing that we need to keep in mind though is that the two matrices needs to have the same size. So I mean that both Matrix A and Matrix B need to have a dimension, so dimension of Matrix A should be equal to dimension of Matrix B, and let's say should be equal to M by N, but for the rest we don't really need to care uh which one we will put first; will we put first A and then add the B, or we will do the other way around. So we will then first take B and then we will add A. So this is the idea behind commutative law for matrix addition. So first let's look into all this uh laws, and then we will also look into the corresponding examples. So for this specific case it might actually also be helpful to write down the general formula which will um make sense out of this um law for the uh uh which is a commutative law for the matrix additions.

So let's say we got a matrix A which is M by N, and this Matrix can be represented as A11 dot dot dot AM1. So this is something that we saw time and time again, so I'll just quickly write it down, the common notation for this, and then here we have the last column which is AMN, and this is the Matrix A. Then we got Matrix B which is again M by N and can be represented as B11 and then dot dot dot and then BM1 dot dot dot B1N dot dot dot BMN. So the commutative law says that A + B should be equal to B + A. Let's check that whether this is the case. Let's first compute this part, and then we will do this one. Well, the first one means that we get, so A + B is, and we remember remember how we add matrices, right? So we know that we just need to pick their corresponding elements and add them to each other. So we get A11 + B11, then dot dot dot and then AM1 + BM1. This is why also D is really important that they got um the same dimension, which means that they got exactly the same amount of elements, the same uh amount of columns and the same amount of rows um in terms of the uh Matrix size. So then here we have A1N and then + B1N, then dot dot dot and then AM1 and then + BM. Here I need to put N; we are in the last element of the Matrix, so BMN, and that is it; this is our Matrix A + B. Let's now look into the Matrix B + A. So what that amount is. So the Matrix B + A will then be equal to B11 + A11 dot dot dot and then BM1 + AM1, then dot dot dot and then B1N + A1N, then dot dot dot and then the last element will be BMN and then + AMN. So in here, if we remember from the real numbers, we know that A + B is equal to B + A. For instance, if A is equal to 2 and then B is equal to 1, then A + B is = to 2 + 1, which is equal to 3, and then B + 1 is = to 1 + 2, and it's equal to 3. So we know that indeed for the real numbers A + B is equal to B + A, and making use of that property we can already state that B11 + A11 is equal to A11 + B11, and then the general case is that AIJ + BIJ is equal to BIJ + AIJ, where I is the index of the rows and then J is the index of the columns from the coefficient labeling. So using this property from the real numbers, given that all these values in the M Matrix are real numbers, we can quickly see that the Matrix B + A that we just got in here is equal to this Matrix A + B, which means that 1 is equal to B, and this proves that A + B is equal to B + A. This is the commutative property of the matrix additions.

So the next law is the associative law for matrices, which says that the in case of matrix addition, A + B + C is equal to A + B + C. So basically this time we go from here to adding one more element, which is the third Matrix, Matrix C. So we are saying A + B + C is equal to A + B + C. So it doesn't matter whether we will first take the Matrix A and then B and then add them up, and then we add Matrix C, or if we first take the Matrix B and C, add them up, and then we add A to this sum; it doesn't matter; we will see an example of this in a bit. And then the uh second part of this associative law for matrices says that for matrix multiplication, A * B * C is equal to A * B * C. So again, in terms of the um order when it comes to this specific multiplication, so it doesn't matter whether we will first multiply A by B and then by C, or we will first multiply B by C and then we add the A at the end; we will end up with the same amount. So A * B and then * C is equal to B * C, and then in the left-hand side we add the A. So A * B * C. So these properties help us to add or multiply matrices without really worrying about this idea of grouping of the terms. So we can always group them and perform all sorts of operations. So this is this first property that we see in here. Let's say we have this uh Matrix A, matrix B, and Matrix C. So let's prove that in the order doesn't matter and this associative property holds. So let's prove that. So for that, the first thing we need to do is to compute A + B, so this part. So A + B + C, what is that? First, I need to compute this part, and then I will add C, which is the second part. So A + B is equal to my A is equal to 1, 2, 3, 4 plus and my B is equal to 5, 6, 7, 8. This is then equal to, so 1 + 5 is = to 6, 2 + 6 is = 8, 3 + 7 is = to 10, and then 4 + 8 is equal to 12. This is my A + B; this is the first part. Now the second part is then to add to this A + B this C; this I can, by the way, also call some Matrix D, so I can say that this is equal to D + C. So let's find out what is this amount. So A + B, or what we're referring as D, is 6, 8, 10, 12; we just calculated it in here. I'm also adding now my Matrix C, which is 9, 10, 11, 12, so 9, 10, 11, 12. What is this amount? It is 6 + 9 is 15, 8 + 10 is 18, 10 + 11 is 21, 12 + 12 is 24. This is my final Matrix. So I have checked that D, A + B + C is equal to 15, 18, 21, and 24. This is the first part. Let's now go ahead and check whether this is equal to the second part, which is this part. So this is one, this is two. So this is then A + B + C, as you can see it in here. Let's now calculate that amount, and like previously we will do it in an order. So first we need to calculate this part, and then the entire thing. So B + C is then equal to the B was 5, 6, 7, 8, 5, 6, 7, 8 plus and the C was 9, 10, 11, 12, 9, 10, 11, 12. What is the much? 5 + 9 is 14, 6 + 10 is 16, 7 + 11 is 18, and 8 + 12 is 20. This is the first amount. Let's refer refer this as a D, or we can even call it by some other letter, let's say K; this is Matrix K. So B + C is K. Then the second part is to take this B + C, so B + C, which we have referred as K, say K, and then we are adding to this the A, and specifically just to ensure that we stay with the same order, I'm saying I will add from the left side the A to this Matrix K, and this obviously means this is equal to, so A + B + C, this is what I'm referring by just uh in a more simpler note A; I'm just using K in here. So this is my B + C, or what I'm referring also as a K, and this amount is equal to what is my A? My A is 1, 2, 3, 4, 1, 2, 3, 4 plus and what is B + C? We just calculated that that's the K, so 14, 16, 18, 20. So 1 + 14 is 15, 2 + 16 is 18, 3 + 18 is 21, 4 + 20 is 24. So we have learned that the A + B + C is this Vector. Now is this Vector equal to the A + B and then + C? Well, here we got this 15, 18, 21, 24, 15, 18, 21, 24. So we have just proved that the first part is equal to the second part, which means that we have proved that indeed the order doesn't matter, and A + B + C is equal to A + B + C. So this calculation confirms that the both sides of this equations they are in indeed equal, and this confirms the associative law for the matrix addition.

So let's now look into the distributive law for matrices, which says that matrix addition and multiplication they satisfy the distributive property, which means that if we have a left distribution, A * (B + C) is equal to A * B + A * C, and then in the right distribution we basically have the Matrix multiplying from the right, from hence the name right distribution, (A + B) * C is equal to A * C + B * C. You might very quickly see and recognize from here that we have very similar, actually exactly uh the same rule, only for real numbers; we know that A * (B + C) is equal to, and then we open the parentheses with, say this is equal to this times this, so AB plus this times this, AC. You can see that we have exactly the same here, only in the capital letters. So in the real numbers we have exactly the same law, so the same we have also for our left distribution when it comes to matrix additional multiplication, and the same we have only with a different order here; you can see the C. So this one is basically uh with the different order; instead of having the Matrix multiplied in the left here we have from the right, and this is similar to the property that (A + B) * C is equal to C * A, which is AC plus C * B, which is BC. An example uh where we will prove that the distributive law for matrices um is indeed true, and I have skipped deliberately the uh example for this one because uh this A * B * C is equal to A * (B * C). So the associative law for matrix multiplication, it's something that you can calculate for yourself using the same A, B, and C matrices, only this includes multiplication of these two matrices, and it's something that we are going to do as part of this example. So instead of doing and redoing this multiplication, I thought that it's great to leave that for you as a practice and instead focus on bit more complex problem like this one, that one way or the other includes the same matrix multiplication. So I need to calculate the A * B in this case, which means that by providing you this example I'm also including what is needed to do the previous example, only it would be a great way to practice the material for yourself. So let's now move into proving the distributive law for matrices. So we got this matrices A, B, and C, and here I'm going to apply matrix multiplication the same as that is needed for the previous uh case, and here what we need to prove is that A * (B + C) is equal to A * B + A * C. So this is the first part; this is the second part. So let's go and calculate them separately. So for the first one we need to calculate A * (B + C), which is then something that we can calculate by first doing the addition. So we will first do the addition of matrices B and C, and then once we are done with that we will then do A * (B + C). So that's the second part. So let's go ahead and do that calculation. So first we got B + C; what is B + C? B is 5, 6, 7, 8, 5, 6, 7, 8 plus and the C is -1, 0, 0, -1. So on the diagonal we got -1 and -1, and then of diagonal lower and upper part we got zero. And what is this Matrix? This is equal to 5 - 1, so 5 + -1 is equal to 4, 6 + 0 is = 6, 7 + 0 is = 7, and then 8 - 1

is = 7. This is our B + C, which we can refer to also as Matrix D. So let's call this D, which means that now we are interested in a * B.

So for this second part, we need to take this Matrix a. So a * D is then equals to: we need to take the Matrix a, which is 1 2 3 4, and we need to multiply it with this Matrix that we just got, because this is the B + C or the D that we were referring to: 4 6 7 7. Okay, so let me remove this part, cuz we are going to need some space for this, and let's do this calculation. This is 2x2 and this is 2x2. I will do the calculations in here. So we need to end up with the Matrix that is also 2x2, because we know 2x2 Matrix times 2x2: we will pick this part, so the number of rows and the number of columns of the second one; this will be our resulting Matrix, which is 2x2. All right.

So for the matrix multiplication, we know that for this element in the place of, so one a, or let's call this Matrix—we don't even actually need to call this anything; we can keep it simple—so let's say that we are in the first row in the First Column. So this is the first row in the First Column. For this, what we need to do is we need to take the first row from the first Matrix, so Matrix a, and then the First Column of the Matrix D, so this one, and we need to do the dot product, which means that we do basically 1 * 4 + 2 * 7.

So for this element, which is in the second row and the First Column, we need to take the second row and First Column in here. So we end up with 3 * 4 + 4 * 7. The dot product between this one and then this one. So for this element, which is in the first row and then the second column of the final Matrix—so first row and second column—we need to pick the first row and second column and do a dot product, which means 1 * 6 + 2 * 7. And then in here, in this element, we got the second column and second row, so second row second column, which means that we need to pick the second row and the second column, the dot product of the second row of Matrix a and the second column of Matrix D, which is 3 * 6 + 4 * 7.

So let's quickly calculate what this amount is. So this is the a * B + C basically, and this Matrix is: 1 + 4 is 4; 2 * 7 is 14; 4 + 14 is 18; 1 * 6 is 6; 2 * 7 is 14; and 6 + 14 is 20; 3 * 4 is 12; 4 * 7 is 28, which means that we got here 40; 3 * 6 is 18; 4 * 7 is 28, which means we got here 46. This is our final a * b + C. Let's now go ahead and calculate the second part. So the second part says that we got a * b + a * C. So a * b + a * C, which means that first we need to do this calculation and then this one, and then we need to add them to each other. So let's quickly then calculate what is a * B and then a * C and then add them to each other. Let me clean up some space in here. We're going to know that when we write this one in a smaller format, so this is equal to 18 20 40 and 46. And let me take over the second element which we still need to calculate, which is aB + aC. First we will do this and then this, and then we will add them to each other. So a * B is equal to 1 2 3 4 multiplied by 5 6 7 8.

Now, following the same approach from the previous example when we calculate this Matrix, I will then quickly calculate what is a * B. So in here we got first row and First Column, so 1 * 5 + 2 * 7, so the dot product between the first row and the First Column from here. Now for this element here, we got the second row and the First Column; we need to take the second row in the First Column from here and we do the dot product, which means 3 * 5 + 4 * 7. In here, here we got the first row and second column, so the first row and second column, which means that we need to have 1 * 6 + 2 * 8. Then here we got the second row and then second column, which means 3 * 6 + 4 * 8. And then this is equal to: 1 * 5 is 5; 2 * 7 is 14, so this is 19; 1 * 6 is 6; 2 * 8 is 16; 6 + 16 is 22; in here 3 * 5 is 15; 4 * 7 is 28, so this is then 43; and in here we got 3 * 6 is 18; 4 * 8 is 32, so we end up with 50. So hope I haven't made any mistakes in the calculations. So this is the a * B, so a * B is then equal to 19 22 43 50. Let's clean this space and let's move ahead to the second part of the calculation, which is a * C. What is a * C? Well, a * C is 1 2 3 4 multiplied by -1 0 0 -1. So here we are then getting -1 + 0; here we are getting -3 + 0; here we have -1 + 0, so no, so the first row and second column which is 0, 2; and then in here we got the second row and the second column which is 0 - 4, which means that we end up with this Matrix and it's equal to -1 -3 then -2 and then -4, which means that we are getting as a final step aB + aC, which means 19 22 43 and then 50, then plus -1 -2 -3 -4, and what is this? 19 - 1 is 18; 43 - 3 is 40; 22 - 2 is 20; 50 - 4 is 46. Okay, so we got that this amount aB + aC is equal to 18 20 40 46. And as you can see already here, this Matrix that we got in the previous calculation from one is equal to this Matrix that we got as part of the second calculation, which means that now we have proved that for this specific example indeed 1 is equal to 2, which means that a * b + C is equal to aB + aC. There we go.

So let's now look into another law, which is the scalar multiplication law for matrices. So the scalar multiplication law for matrices says that if we got a scalar R and a matrix A and B, then R * a * B is equal to R * a * B and is equal to a * R * B. So here the r is just a scalar, so it's a real number, and then A and B are matrices. And what this law basically says is is that it doesn't matter in which stage you will do your matrix multiplication with the scalar. If you have this external scalar, you can first take the two matrices, multiply them with each other, so the A and then B, and then multiply it with r, or you can take the scalar R, multiply with your first Matrix, and then multiply with B, or you can take your second Matrix, multiply with the scalar, and then multiply with a. It doesn't matter; they will all result in the same Matrix. So let us actually prove this by making use of our skills from matrix multiplication and scalar multiplication. Here I've picked up a bit more advanced example where a is 2x3 and B is 3x3. In this way, we will train our multiplication skills for matrices, and at the same time we will also prove that the scalar multiplication law of matrices holds. So let's go ahead and do the multiplications. So first we have a matrix a. What is that Matrix? Matrix a is 1 -1 2, so 1 -1 and then 2, then we got 0 2 and then 1, which is 2x3, and then we got B which is equal to: it is 3x3 with elements 1 0 1, 1 2 0, 1 1 3, 1 0 2. So it is 3x4, so it's 3x4, not 3x3, but 3x4 Matrix. Now the final part that I need here is this, which is R is equal to 2, the scalar value. So R is equal to 2. So the first thing that I'm going to do is to calculate this amount, which is R * a * B. For that, what I need to do is to first calculate this a * B. So let me quickly go and calculate this for us. So given that the a has dimension 2x3 and then B has a dimension 3x4, I can see that quickly that my dimension criteria is satisfied: the number of columns of a is equal to the number of rows of B, so that's fine, and then I know also know that the final dimension of a * B is going to be 2x4, so it's going to be a 2x4 Matrix. And how do I know that? Well, because I know that from our um all the problems that we have solved, we have already seen that we always need to pick the number of rows of the First Column and the number of columns of the second uh Matrix in order to get the final Dimension, which is 2x4. So let me then go ahead and do the calculation. So we are going to have a 2x4 Matrix. Let me write it even bigger, so it's going to be a 2x4 Matrix. So for the first row and First Column, I need to look in here, the first row and the First Column, which means I need to take one, so it's equal to 1 * 1 + 1 * 1 is 1 - 1 * 2 is -2; 2 * 3 is 6 + 6. This is my first value, and what is this amount? It is equal to 1 - 2 is -1 and 6 - 1 is equal to 5. So this amount is five, five. So what is this amount? Well, this is my second row in the First Column, so I need to make use of second row and First Column, which is equal to: 0 * 1 is 0; 2 * 2 is 4; and 1 * 3 is 3; 4 + 3 is 7. So this value is 7, seven. We are ready to go on to the next column, so column number two. So then this time I need to look at the first row and second column. So we are going to use this one, so first we will use this first row: 1 * 0 is 0; -1 * 0 is 0; 0 + 0 is 0; and then 2 * 1 is the only nonzero element; 2 * 1 is two. So I already know that for my second column I got here two. And what is this element? Well, for this I need to look at the second row and second column, so this thing: 0 * 0 is 0; 2 * 0 is 0; 1 * 1 is one, which means that here I get a one. Let's now move on to on uh towards the third column. So in here first I need to look at the first row, so 1 -1 and 2, and then this time remove this; I need to look at the third column because I'm here in the third column: 1 * 1 is 1; -1 * 1 is 1; so here I got 1 - 1 and then 2 * 0 is 0 + 0; 1 - 1 + 0 is 0, because those two cancel out. This means that here in this element I got a zero. And what about this element? Where I need to look here in the second row and here I need to look at the third column: 0 * 1 is 0; 2 * 1 is 2; 1 * 0 is 0; so 0 + 2 + 0 is equal to 2. So this is 2. And now we are left with the fourth column. So for that I need to look in here. So for the first row, which means first row in here and then the fourth column in here, so first row in here and fourth column here: 1 * 1 is 1; -1 * 1 is 1; and then 2 * 2 is 4, which means that I end up with 1 - 1 and then + 4. And what is this? This two cancel out; I end up with four, which means that here I need to fill in four. And what is this final element? It is the second row in the fourth column, so the second row in the fourth column: 0 * 1 is 0; 2 * 1 is 2; 1 * 2 is 2; 0 + 2 + 2 is = 4. So now we obtained that a * B is this 2x4 Matrix as we have expected. So this is then equal to 5 2 0 4 and then 7 1 2 4. So then the next step would be to take the scalar R and multiply it with a * B. Let me actually keep the colors consistent, so a * B, this is a * B. So the only thing that I need to do is to take that in here and multiply this two with each of those elements. So I will end up with the same size Matrix, so 2x4, only all these elements need to be multiplied with the scalar, which means that I will get: 5 * 2 is 10; 2 * 2 is 4; 0 * 2 is 0; 4 * 2 is 8; and then 7 * 2 is 14; 1 * 2 is 2; 2 * 2 is 4; and then 4 * 2 is 8. So this is the result of the multiplication. So this is the first part; this is what we are referring to as one. So we have then checked in here that the r times—actually we have already in here—so I won't be writing again. So as part of the first section we have already seen that R * a * B is this Matrix. Let's now move on to the next one, which is calculating the second part. So this is the first part, this is the second, and this is the third. We have this already. Let's now move on and calculate this one. So for this second case, so the second case what we want to calculate is R * a * B, so it is R * a and then * B. This is what we need to calculate. So the first thing that we will do is to calculate this part and then to calculate the entire thing, the second point. So let's go ahead and do that. First we will take the A and then we will multiply all its elements by scalar two to get the r and then a. This amount is equal to: 1 * 2 is = 2; -1 * 2 is -2; 2 * 2 is = 4; 0 * 2 is = 0; 2 * 2 is = 4; 1 * 2 is = 2. This is R * a. Now in the next step, so this was one, the next step we need to take this amount, this Matrix 2 -2 4 0 4 2, and multiply it with 1 0 1 1 2 0 1 1 3 1 0 2; so basically the Matrix B. Let's now move and work our way out with that one. Actually, let me remove this from here and keep the space a bit more clean: R times a, and I will be multiplying this with the Matrix 1 0 1 1 2 0 1 1 3 1 0 2. Well, I know that this one is 2x3 and this one is 3x4, which means that the result will be 2x4. Let's now go ahead and calculate that Matrix, which is equal with a dimension of 2x4. Well, for the first row and First Column, let me actually go and quickly do those calculations. Let's now go ahead and do those calculations. So we are going to have four columns as previously; the dimension is going to be 2x4. So let's do it column by column. In here it means that we are in the row one and then column one: 2 * 1 is equal to 2; -2 * 2 is -4; and then here we got 4, so 4 * 3 is 12; so we got 2 - 4 and then + 12; and what is this amount? 2 - 4 is -2 + 12 is 10. So here we got 10. Let me remove this 10. This is the second row and the First Column, which means we got 0 * 1 is 0; 4 * 2 is 8; and 2 * 3 is 6; so 8 + 6 is equal to 14. So here we got 14. This is the first row and second column, which means that we are looking at this row and second column this time: 2 * 0 is 0; -2 * 0 is 0; the only thing that we care about is this one and this element, which is 4 * 1, so this should be four. Let's now do the same for the second row: 0 * 0 is 0; 4 * 0 is 0; 0 + 0 is 0, which means we are left with 2 * 1, so here it comes two. Let's now do the third column. So for the third column we got First Row: 2 * 1 is 2; -2 * 1 is -2; and then 4 * 0 is 0, which means that here we get 0 because 2 - 2 + 0 is 0. Then we got the second row and third column, which is this row and then third column: 0 * 1 is 0; 4 * 1 is 4; 2 * 0 is 0; 0 + 4 + 0 is four, so this is four. And then for the first row and then fourth column, so it means that we need to look at this specific column: the first row is 2 * 1, it is 2; -2 * 1 is -2; and then 4 * 2 is 8; so 2 - 2 + 8 is 8. And then finally for the second row and the fourth column: 0 * 1 is 0; 4 * 1 is 4; 2 * 2 is 4; 4 + 4 is 8. This is the final Matrix, which means that this entire amount that we just calculated step by step, this is equal to this Matrix in here. Okay, so this is the second element. Let's check whether the first element is equal to the first one. So we see here 10 4 0 8 14 2 2 8. As you can see, we are dealing with exactly the same Matrix, which proves that indeed R * a * B is equal to R * a * B. So this part we have already proven because we have seen that 1 is equal to 2. Perfect. So the only thing that is remaining is to calculate this third part and to see whether this is equal to this matrices, because we have seen that the two of those are equal. So the remaining thing that is left to prove this theorem is to calculate this third part. Let's go ahead and do that. So the third element says: let's first calculate the r * B and then multiply it by a. So we need to calculate r * B * a. This is what we need to calculate, which means first we need to calculate this and then we need to calculate the entire thing. All right. So let's go ahead and do that. So R * B is equal to: so we need to multiply each of the elements of B by two, so we end up with this Matrix: 2 0 2 2 and then 2 * 2 is 4, 2 * 0 is 0 and then 2 2 and then 2 * 3 is 6, 2 * 1 is 2, 2 * 0 is 0, 2 * 2 is 4. This is that first Matrix. Let's now go ahead and calculate the second part, which is a * Matrix a, so it is 1 -1 2 and then 0 2 and then 1 multiplied by this Matrix which is 2 4 6 0 0 2 and then 2 2 0 and then 2 2 4. Okay, perfect. So this is then what we need to calculate. Well, this is 3x4; this is 2x3, which means the result should be 2x4. Let's go ahead and do those calculations. This first amount...

Will be the first row and the first column. Dot product of those, which means 1 * 2 is 2; -1 * 4 is -4; and then 2 * 6 is 12. So here we got 2 - 4 + 12. And what is this amount? Well, 2 - 4 is -2; 12 - 2 is = to 10. So this one, this element is 10. Then for the second row, we need to look in here: so 0, 2, 1, and the dot product of the row with the first column, so this thing and that is 0 * 2 is 0; 2 * 4 is 8; and then 1 * 6 is 6. So what is 8 + 6? That is 14. And then for the first row and then the second column, we need to look to the first row in here and then the second column in here and the dot product of the two. Well, 1 * 0 is 0; -1 * 0 is 0; and then 2 * 2 is four. So that's what we are left with: four. And then for the second row and then the second column, so this element, we got 0 * 0 is 0; 2 * 0 is 0; 1 * 2 is 2.

For the first row and the third column, so we need to look in here: 1 * 2 is 2; -1 * 2 is -2; and then 2 * 0 is 0. So we are left with zero. 0. And then once we do the calculation for the second row, we will see that we end up with 0 * 2 is 0; 2 * 2 is 4; and then 1 * 0 is 0. So we end up with four. And then here for the final column: 1 * 2 is 2; -1 * 2 is -2; and then 2 * 4 is 8. The first two cancel out, and we end up with 8. And then for the second row and the fourth column, we look into here again, this time the second row: 0 * 2 is 0; 2 * 2 is 4; and 1 * 4 is 4; and then 4 + 4 is 8. So if we look in here, this is our third amount. We will quickly see that again we have the same matrix with exactly the same elements. So now we have also proved this third part, and we have seen that in all cases the R * A * B is equal to R * A * B is equal to A * R * B.

Now we are ready to move on towards the second module in this unit, which is about the determinants and their properties. We are going to look into the determinants at a high level; we are going to define them and going to understand what, why they matter, and why they are important. Then we are going to see how we can calculate the determinants. We are going to see the calculation for a 2x2 matrix, then a 3x3 matrix, and then just in general how we can do it. And then we are going to see the properties of determinants one by one, and then finally we are going to see the determinants interpretation from the geometric perspective, so when we visualize it using Python.

By definition, the determinant is a scalar value that can be computed from the elements of a square matrix. So this is important: square matrix, and encodes certain properties of the matrix. So the determinant provides critical information about the matrix, such as whether it's invertible and the volume scaling factor for the linear transformation it represents. So we see that the concept of the determinant is highly related to many other concepts that we have seen before. So first here it's talking about the square matrix; then it's talking about encoding certain properties; so having the determinant, it contains certain information that is related to the properties of the system that that matrix is representing; and then it provides critical information about the underlying matrix. Because the determinant is calculated from a matrix, we say the determinant of a matrix, so it contains critical information about that matrix, such as whether it's invertible or not. And this goes back to the concept of inverse; we will see this once we learn the concept of determinant because the inverse calculation is dependent on the determinant. But keep this thing in mind that the determinant contains information whether we can get an inverse from a matrix or not. We will see this concept also in detail in the next section, but for now we can remember that the determinant contains important information related to the invertibility of the matrix: so having an inverse or not. And then it also contains information about the volume scaling factor for the linear transformation it represents. So here we then go back to this concept of Ax = B, and then knowing the determinant, we can then comment on this volume scaling factor for this linear transformation that it represents. So let's go on to the next slide to find out a bit more about the determinants and specifically how we can calculate the determinant in mathematical terms when it comes to the 2x2 matrix.

Because the determinant of a 2x2 matrix is quite straightforward. So for a 2x2 matrix A with these elements, where a, b, c, and d they are all real numbers, the determinant, which we define by this det(A) – so that is a short way of saying determinant, and then in here we always write the matrix of which we are computing the determinant – is then equal to, and then we are taking this a * d, so we are taking this diagonal elements a * d, so they are on the diagonal, and then we are subtracting from this this other two, the remaining two elements of the diagonal, so b * c. And this gives us the determinant of a 2x2 matrix. This is just a formula that you need to remember whenever you want to calculate the determinant of a matrix by hand, manually. So the calculation for larger matrices involves a bit more difficult calculation. We will see also in a bit the determinant of a 3x3 matrix. It relies on the determinant of a 2x2 matrix, and the idea is that every time we increase the dimension of our problem, so let's say we are in R4, then we will go back to the R3, and then given that R3 relies on the determinant of the underlying 2x2 matrices, anytime we increase the dimension, we again go back to this idea of using 2x2 matrices that form the entire matrix in order to compute the determinant, only when it is R4, R5, etc. So it becomes much more difficult to describe and to do it manually; therefore, there are other algorithms which we will see at the end of this course, like decomposition algorithms and factorization algorithms that can be used in order to calculate the determinant of a matrix that has higher dimension, higher than three, for instance. But in this specific unit, we are going to discuss both the calculation of the 2x2 matrices determinant and the determinant of a 3x3 matrices, and we will also see detailed examples of them. So without further ado, let's then go ahead and calculate the determinant of this 2x2 matrix.

Let's now look into this specific example where we are calculating the determinant of this 2x2 matrix. So this is the A, and let's keep in mind that this is the a, this is the b, in this, not in this way of writing the matrix A, so the letters corresponding to the elements of this matrix A, so this is the a, this is the b, and then this is the c, this is the d. And we said that the determinant of A is equal to the diagonal elements, so 1 * 4 minus the off-diagonal elements, which is 2 * 3, because we said that the definition of this determinant is that is equal to a * d and then minus b * c, which is exactly what we are doing in here. So if we calculate 1 * 4 is = to 4, and then 2 * 3 is = 6, 4 - 6 is = to -2. Therefore, we say that the determinant of matrix A is equal to -2. Let's now go ahead and practice with calculation of determinants on two other matrices. So in this case we are still in the two-dimensional space, so we have 2x2 matrices. We'll first calculate the determinant for matrix A. So we see that we got this element 5, 0, 6, 1, and we know that by definition the determinant of the 2x2 matrix, so that of matrix A is equal to a * d - b * c, where the matrix has the following form: so we got a and then d in here, and then b and the c in here. So we see that this is basically our a, this is our d, this is our b, and this our c. The way you can also say is that those are the diagonal elements and those are the off-diagonal elements. So therefore it means that we can calculate the determinant of matrix A by taking the five multiplying with one, so it is 5 * 1 minus the off-diagonal element which is 6 * 0, and this amount is equal to 5 - 0 and is equal to 5. Let's go ahead and also calculate the determinant of matrix B. We see here that on the diagonal we have this two elements 1, 1, and off-diagonal elements are both zero, those two. Therefore, we can calculate the determinant of this 2x2 matrix, which is also sometimes referred as I2, so it is the identity matrix because we got here the E1 and then E2 in the two-dimensional space. And the determinant of the matrix B using this definition is then equal to 1 * 1 - 0 * 0, so 1 * 1 - 0 * 0, and this is equal to 1. And this is actually a special case of determinant, and later on we will see why it is so important to have this relationship of identity matrix having a determinant and having it equal to one, and this relationship between determinant and identity matrix is something that we see in the upcoming lesson, so keep this one in mind.

So now when we are clear on how we can calculate the determinant for a 2x2 matrix – so this is quite simple and straightforward calculation by taking the diagonal elements a and then d and then subtracting from that, from that product a * d, we are subtracting the off-diagonal elements products b * c – we can then get our determinant. And now when we are clear on that, we are ready to go on to a bit more advanced calculations, which is calculating the determinant this time for the 3x3 matrix. So now we increase the dimension size and we go from R2 to R3 because now we have a 3x3 matrix. And by definition, given a 3x3 matrix A which has the following elements: so a11, a21, a31, and then a12, a32, a32, so we have already seen this coefficient labeling, this should look very familiar, this is a 3x3 matrix, and the determinant of a matrix A, denoted as det(A), is calculated using the formula. And here we see the formula; we are basically using the 2x2 matrices that form this matrix A in order to calculate the determinant of the 3x3 matrix. And how we are doing that? Well, we are using this element and then this element and this element, and every time we are hiding part of the matrix. So when we have, for instance, this a11, so for this first part we are saying, well, let's hide the row and the column corresponding to this element, which means that we need to hide this row and this column, and what is left is this 2x2 matrix. We will calculate the determinant of this 2x2 matrix, and we will multiply this with this element that we use in order to remove the corresponding row and column. This will form the first element in here. So you can see a11, which is a simple value, so this is the entry value which is in the first row and first column, a11, multiplied by the determinant of this matrix, so this matrix. So once we have that and we already know how we can calculate a determinant of a 2x2 matrix because this 2x2, so taking the diagonal elements and then multiplying them together, subtracting from that the off-diagonal elements product, now we are ready to go on to the next part of the calculation, which is this time adding a minus here. So you can see here this here is plus and then here is minus, so we do here minus. And for this second step, what we need to do is kind of similar, only this time the element that we will be using to understand how we can remove the row and the column, so we will then dark it out, it is this one, a12. So then we will need to remove this column and this row, and then the remaining matrix, which is this one, this 2x2, and here I mean a21, a31, and then a23, and then a33, this is the matrix that you can see in here remaining, which means remove this one and then this one, and then the remaining 2x2 matrix is what you need to use in order to do your calculations. So you can see that I got exactly the same in here, and once again we are computing the determinant of this matrix; we are multiplying this with this a12 element, so this element, and now we have also the second element in our calculation. And then we go on to the next step, which is a plus sign here; let me use the same colors, plus sign here, and then we are using this time our final third element to understand which row and which column we need to dark out, which is this element. So we then remove the first row and the last column, and this is then the matrix, the 2x2 matrix that we use in order to do our calculation. So det(this matrix) multiplied by the a13. So we could also use in the same manner this row or this row; it really depends on the kind of values. The tip that I will provide to you, or the trick, is that to always look for these zero values. Wherever I see zeros or I see ones, I'm thinking that, hey, those zero values that give me the more straightforward and easy calculations because if I have zeros in my calculation, in my entry, so if I got a zero here, for instance, 0 times any determinant is zero; I don't even then need to calculate the determinant, right? Because then I know that I'm multiplying that determinant with zero. Therefore, if I know that that entry, for instance, this row contains the majority of zero, so it is 0, 0, 1, then of course it's a great row to pick to use these target elements. So in that way, I will then know that this is the row that I need to target. But if it is like that, that for instance I got a matrix 1, 0, 3, 4, and here I got 0, 1, 0, and here I have 1, 0, 0, 3, and 4, of course the easiest thing would be to not use this row but instead use this one. So in that case, I will then have this zero and zero as my target values, which means that I will only need to calculate the determinant of a 2x2 matrix. This for this one, for the two cases, I don't need to do it because I know that the corresponding target values, the target elements from my matrix will be zero. So let me show you what I mean by that. So if, for instance, I go for this second row and not the first one, what I need to do is that I can calculate the determinant of A by taking the a21. This then will be my target; I will then need to remove this column and this row; then I will need to do the determinant of a12, a13, a32, a33. This is what then I need to do. Then the next thing I need to do is, of course, here I have a plus, here I need to do minus because we always need to interchange the values, so here is a plus, here is a minus, here is a plus. So I do plus, a minus in here, then I do the next element in my row, which is this one; let me use red color, so a22, so I'm then doing a22 multiplied by the determinant of, so I'm removing this row and this column, a11, a13, and then a31, a33. And then the final part is, of course, as you might have already guessed, is to look into this element, so it is plus a23 multiplied by the determinant of, let me actually write it down in here, the determinant of a11, a12, and then a31 and then a32. So in this way, basically, independent of what row I will take as my leading row that I will do my calculations, and I will just need to pick one row; I can always get the same value for determinant of A. But choosing intelligently which row to pick will save you a lot of time and headache in terms of calculations because if you are dealing with a row that contains many zeros, for instance, you have 0, 0, 1, or 1, 0, 0, or even better 0, 0, 0, then you know automatically that you will need to calculate your det, the determinant once here, also once, and here you don't even need to calculate it; you know that you got zero here, zero here, zero here, so it's automatically equal to zero. So I hope this makes sense because this is a trick that usually you will not come across, but this just helps you to save a lot of time when it comes to calculation of your determinants in a 3x3 setting.

In this case, we have this, we now we have this definition and we know the tricks that we can use, but I think it's really helpful to go ahead and to solve a problem. So basically, this is the higher-level summary of the steps that we just discussed. So the determinant of a 3x3 matrix simply involves multiplying the a11 by the determinant of the 2x2 matrix that remains after excluding the row and column of a11, so what we did in here by doing this, and subtracting the product of a12 and the determinant of its respective 2x2 matrix, so this part, and then adding the product of a13 and the determinant of its respective 2x2 matrix, so this part, and the signs alternate, so it means first you always got the plus, then you always get the minus, and then the plus, so they interchange; you start with plus, then you do the minus, and then the plus. So let's go ahead and calculate the determinant of this matrix. So before even looking at the answer, let's actually go ahead and do that on this page, paper. So we got a matrix A which is equal to 1, 2, 3, 4, so basically from 1 till 9, 1, 2, 3, 4, 6, 4, 5, 6, and then 7, 8, 9. And for this 3x3 matrix, we need to calculate the determinant. So the determinant of A, the first thing that I'm seeing is that there are no rows with zeros or columns, which means that I cannot use my trick, and instead I will just need to go with, let's say, the first row, and it's also convenient given that I got as scalar these values, these much smaller values relatively to the other ones. All right, so first things first, let's go ahead and write down that formula. So the determinant of A is equal to: first we are going to take this one, so our 1, 1 times, and then we got determinant of, and then we have this matrix which is 5, 6, 8, and 9. This is our remaining matrix. Then the next thing we need to do is to interchange the signs, the sign, which is minus, and then we got, so the remaining matrix is then determinant of 4, 6, 7, 9. And then finally, plus, plus three times, and then determinant of what do we have? Well, this is the target, so it is 4, 5, 7, 8. Right, so let's go and do those calculations quickly. This is equal to 1 * the determinant of this is the diagonal element, so 5 * it is 5 * 9 - 8 * 6 - 2 * 4 * 6 - 7 * 6. Sorry, 4 * 9, so the diagonal elements 4 * 9 - 7 * 6. And then plus three times, and then 4 * 8 - 5 * 7. This is equal to: so 9 * 5 is = 45; 8 * 6 is = 48; then -2 * 4 * 9 is 36; 7 * 6 is = 42; + 3 * 4 * 8 is = 32 - 5 * 7 is = 35; so this...

is equal to 1 * -3 - 2 * and then here we got 36 - 42, so that's -6 + 3 * -3, which is that = to -3 + 12 - 9, which which is equal to zero. So let's check it; indeed we got the right answer, perfect.

So now when we are clear on how we can do this calculation, let's now go ahead and calculate yet another determinant of a 3x3 Matrix. And this time I want to show you this uh simplified version by you making use of this trick that I uh specified. So instead of using this first row as an indicator, I will be using the uh second row as my indicator. One thing to keep in mind when making use of this trick is that when you start from the second row, so from the even rows—even rows, second, fourth, or sixth—then in those cases you need to flip the signs that you will be using. So while in here you had plus sign plus in the beginning, then you got a minus and then a plus when doing all these calculations, so you remember here we got plus minus plus; when you start from the second row instead of first one, you need to flip the order of this, so you need to start with minus; you have minus, you got plus and then minus.

So knowing this trick, it also means that you go one step beyond and you know how you need to intelligently uh reduce the time that you are spending on calculation calculation of the determinant, but it also means that you need to be careful on knowing what kind of signs you need to use because if you start from the first row you start with plus and then you do minus plus minus plus, so knowing how to start you already know how you can go on, but when it comes to the second row, so the even rows, you need to start with a minus, so you need to do minus plus minus plus dot dot dot. All right, so let's now go ahead and use that technique in here. So here I see that my first row doesn't contain zeros, but my second row does; so this gives me indication that I can reduce the time that I spent on calculating the determinant at least one time because I didn't no longer need to calculate that determinant. So the determinant of B is then equals you; I will then start with minus, given that I'm going to use this row, and then I have zero times—so because this is my element—the determinant the determinant of 2306, and then plus, then this time the second element target element is this one, so it's four times and then we got determinant of 1 3 1 6, and then minus the five times, so this five determinant of 1 2 1 0. And what is this amount? It is equal to this; I don't need to calculate because I got a zero in here. This this trick is all about this—to not calculate the determinant too often. And then this equals you four times four times determinant of this is 6 - 3, so 1 * 6 - 1 * 3, which is equal to 3, and then minus 5 * determinant of 1 2 1 0, which is 0 - 2, so this equal to 4 * 3 - 5 * -2, which is equal to 12 + 10, and this is equal to 22. Let's go ahead and check this, and this is the more detailed and formal derivation.

So one uh interesting thing is that I calculated with my second row, and in here in this slides you can see calculation with the first row. This is just a nice way of seeing the difference that you can do, and here uh in this solution what we have is that we have manually calculated this first determinant too, so in total three determinants, but we again in end up with the same determinant. So independent what kind of row you will use in order to calculate your uh determinant of Matrix B, you will all always end up with the um with the same similar volume unless you have made a mistake in your calculations. So you just need to keep track of the uh rows that contain many zeros and you need to um be careful in terms of the signs that you need to use and the sign that you will need to start. If you start with the first row then start with plus; if you start with the second row then it is minus and then plus Etc. So as you can see here it's a plus and then minus and then plus. In my case I did with my second row, therefore I started with minus. All right, so let's now move on to the properties of determinants.

So the determinant of an identity Matrix is one. That's something that we have also seen when doing our calculations because we are so that in one specific case when we had this example, so this Matrix B and the Matrix B was the identity 2 in the two dimensional space, we have calculated its determinant and we saw that it's equal to one, and this was not a coincidence because the determinant of identity matrices is always equal to one. Then the second property is that swapping two rows or columns of a matrix changes the sign of its determinant. So if you swap rows or columns in your Matrix, so if you end up with Matrix A and B, they are exactly the same only one swaps the two columns or two rows, then you are changing the determinant of that Matrix uh the sign of that determinant but not the value itself. It means that if you got a and you got B and your a is equal to let's say uh A1 and then uh A2 and then A3, so it contains these columns, and then Matrix B is equal to um let's say A2 and then A1 A3, then the determinant determinant of a will be equal to minus of the determinant of B. You can also say determinant of B will then be equal to the minus determinant of a. This is basically the idea of this property. Let's now move on to the third property which says that if a matrix has a row or a column of zeros its determinant is zero. So if you got a matrix a that contains this different values a11 H1 dot dot dot a and uh M1 and then here you got suddenly um column that contains all zeros and then the rest are nonzero even, so in that case you know that your determinant is equal to zero. So for a specific example if you got for instance Matrix 1 2 0 0 0 3 13, then the determinant of this Matrix is equal to zero. And otherwise if you got a matrix B that has a column of zeros, so column that is entirely of zeros, so let's say here we have 1 1 1 and we have a zero Vector here, so we got in here 0 0 0 and then 3 4 5, then given that we have here this zero Vector, then the determinant of Matrix B is equal to zero. And this actually straightforward to be seen from this calculations that we saw because if you do the uh if you pick this specific row and then you do zero times the ter DET minant of the remaining Matrix, 0 times determinant of the other Matrix and then plus, so PL and then so minus and then plus and then minus 0 * determinant of the third Matrix, it is obvious that 0 * a determinant is 0, 0 * determinant is 0, 0 * determinant is zero, which means that you got a whole bunch of zeros to be added to each other or subtracted from each other. This means that if you have a row or a Col with zeros, this already gives you an idea that your determinant is equal to zero; you don't even need to do calculations.

So the final property of determinants is that if a determinant of a product of matrices equals the product of their determinants, so the determinant determinant of a b is equal to determinant of a multip by determinant of B. This is basically what this property is about. So let's quickly go through examples to ensure that we are at the same page with all these properties and we can prove them. So let's say we have an identity Matrix n by n which is we are dealing with I in. Now according to this first property when we calculate the determinant of this Matrix, so determinant of i n is equal to 1. Let's actually look at a specific example, so here we got um identity um Matrix in the two dimensional space in the R2 and we can quickly calculate the determinant of this I2 and we can see that it is equal to this diagonal element, so 1 by 1 - 0 * U; it's actually something that we did as part of my previous examples, so this equal to 1 - Z and is equal to 1. One thing that I wanted to show you before moving on to the next example about the swapping rows is that when we are swapping some of the rows or some of the comms or two rows or two cars of Matrix a, we are referring to this matrix by this notation, so we add this nod in here and we say that that this is basically the manipulated version of Matrix a. So if we have for instance Matrix a equals u a b and and c and those are vectors and then we are swapping two of the Columns, let's say we are swapping this two, we get B and then a and then C, then this Matrix will are referring as a not. This is just an a matter of notation and we just learned that as part of the properties that the determinant of this new Matrix is equal to minus the determinant of a. So if a matrix a has a row or column of zeros then the determinant of it is zero. So let's actually quickly look at this specific example example in here. We got a which is uh having a column of zeros and another column of B and D where B and D are real numbers. So let's prove that this determinant is actually equal to zero. So the determinant of a 2x2 Matrix we have already seen is equal to the diagonal elements, so 0 * D minus the of diagonal Elements which is 0 * B 0 * B and what is number * 0 is equal to 0, 0 - 0 * B is also Z, it's equal to Z, therefore the determinant of a is equal to Z.

So when it comes to the uh determinant of a product of a matrices, let's prove that the determinant of a * B is equal to the determinant of a times the determinant of B. So therefore the first thing we need to do is to calculate this a * B. Let's quickly go ahead and do that. So let me add here this um blank file. So a is equal to 1 2 3 and 4, B is equal to 5 6 7 8, and I I want to prove that the determinant of a b is equal to determinant of a Time determinant of B. First I will be calculating this and then I will be calculating this. So for the first one what I need to do is that first I need to calculate d a * B which is equal to 1 2 3 4 times 5 6 7 8 and then this is equal 2 should be 2 by 2, so first I take this 1 * 5 is 5, 2 * 7 is 14, 14 + 5 is 19, then for this one I need to pick this row, so 3 * 5 is 15 and then 4 * 7 is 28, so 15 + 28 is so there we have 33 43, so I got here 43, then I'm going on to the next column which is in this case 1 * 6 is 6, 6 + 6 16 is 22 and now the second column 3 * 6 is 18, 4 * 8 is 32 and this gives me 50. All right, so now I have the a * B, then as the next step what I need to do is to calculate the determinant of this a * B which is equal to the determinant of this Matrix 19 22 43 50 that I just calculated, and what is this amount? The diagonal elements 19 * 50 - 43 * 22, 19 * 50 is then equal to 950 and 43 * 22 is 946, which means that we end up with four. This means that the determinant of the a * B is equal to 4. Let's quickly check what are the parts of the second amount. So for that I need to calculate determinant of a which is equal to 1 * 4 - 2 * 3, 1 * 4 is 4, 3 * 2 is 6, so 4 - 6 is = -2, determinant of B is equal to 5 * 8 which is equal to 40 and then 7 * 6 is equal to 42 and this is equal to -2 and determinant of a * determinant of B is equal to -2 * -2 which is equal to 4, so we can see that now we just prove that the determinant of a * B is equal to 4. So we have seen that determinant of a is equal to 4 and we see that that's exactly the same as determinant a * determinant of B which is equal to 4, so we have just proven that the this equation indeed holds.

So the determinants they are not just um some calculations or some amounts but they are actually uh important concept and their interpretation um is highly relevant from geometric perspective. So the determinant have a geometric interpretation and the for example the terent of a 2x2 uh Matrix or 3x3 Matrix they represent the area in case of 2x2 or the volume in case of 3x3 Matrix uh of the parallelogram that they are forming. So uh this is often referred as a parallel uh piped um I hope I'm pronouncing this correctly and it's formed by the con vectors of the Matrix. So if we have for instance this uh Matrix a and then we have a b and then C and D, we have this A and C which is the first vector and then B and D which is the second vector and the uh the shoe vectors they actually form a parallelogram um when it comes to the uh two dimensional space and the area that this uh parallelogram um is forming that is equal to the determinant of this Matrix. So the determinant of this scalar value that summarizes this linear transformation that we describe by this Matrix because we saw that we had this a x is equal to B linear system that we were describing using this coefficient Matrix and this was our unknowns, this was our variable and then this B was the um amount that we were uh putting this as equal to. If B was equal to zero then we were solving the homogeneous system, otherwise we had this non-homogeneous system and in the geometric terms the determinant of this Matrix a, so the determinant of a um in case of 2x2 space, so in R2 um when we got two vectors basically in our Matrix a, this is equal to the area that is spent by these vectors in the two dimensional space. In a bit I will also show you specific example such that um we will be on the same page when it comes to this concept of parallelogram, the determinant and those vectors that form the column um uh space of the uh Matrix a. Uh when it comes to the three dimensional space when we have R3, so we got 3x3 Matrix of a, then the determinant of this Matrix a is the volume that is um formed by these uh three-dimensional vectors because unlike the 2D in R3 we got the three vectors that form the a let's say this one this one and then this one and then here we can create this area covered by this three vectors and the area that is formed by the three vectors from a it is equal to the determinant of that Matrix a. So in terms of the 3D it's bit harder to uh visualize it, but in uh case of the two-dimensional space I think this will help uh to improve our understanding of the determinants and make this interpretation uh from geometry uh from geometrical perspective.

So given the two vectors A and B in the two dimensional space, the determinant of this Matrix uh is then equal to the um diagonal elements we already know minus the of diagonal elements, right? So we are also saying we have seen this notation already very often, you will see this volume, this is the absolute; we already know this from high school, this is the absolute volume because the determinant can also be a negative number, we have seen minus 20 or minus 2 and we know that the area cannot be a negative number therefore we are adding this absolute term here. So knowing for example that we have this m matx a which consists of the elements 3 2 and then 1 14, we know that the determinant of this a is equal to 3 * 4 12 - 1 * 2 it is 10 and the absolute value of it, so absolute value of 10 is equal to 10 given that is positive and this is exactly what we have here and this is referred as the area of the parallelogram that the two vectors are forming. And how does that look like in uh the uh coordinate space? So this is the parallelogram that we were referring by and this area that is formed by this parallelogram is equal to the determinant of the a The Matrix a. So so one thing that we need to keep in mind is the definition of parallelogram which means that those two are parallel and they are the same, so this and this lines, those two are the same and then of course the same holes for those two, they are parallel and they have the same um length, therefore this figure in here, this is what we are referring as parallelogram and those two vectors that we can see in here, this one and this one, they form this parallelogram and they are the two vectors that are part of the Matrix a. Hence if we got two vectors that the uh that come from The Matrix a, so Matrix a and we got here this two vectors in a 2x2 Matrix, then the determinant of this Matrix is then describing the area that these two vectors are using or are spanning when creating this parallelogram. So the determinants they play an important role in understanding the geometric properties of the spaces that uh spent uh by these vectors. They provide valuable insights when it comes to the scaling effect effect of linear transformation, the or orientation and the um the locations of them in the cordan system as well as the practical applications in calculating areas in calculating volumes.

Welcome to another unit in our fundamentals to linear algebra course where we are going to talk about Advanced linear algebra Concepts. So uh in the first module we are going to talk about Vector spaces and the projections. We are going to define the bases in a couple of examples of them. We have already touched upon this concept briefly when we are calculating the basis of a no space and the basis of a comp space. We are going to do a similar example in this case and then we are going to uh look into this concept of the uh standard bases for uh different spaces including the R2. We're going to introduce the concept of projections, what is the definition of projections, what is a Formula, how we can calculate it, we are going to look into detailed examples of that. Then we are going to talk about the concept of orthonormal basis. In this module we are going to introduce this concept and we are going to understand the orog gonality normalization. We are going to then discuss a very important topic in linear algebra which is a gramme process. We're going to Define it, we are going to see the overview, the step-by-step process of applying grme uh algorithm, then we are going to see an example of it and the calculations step by step and then we are going to talk about applications of orthonormal bases, the application of gram Smiths process and the importance of this orthonormal basis. This is the module one of this part, so let's first Define the basis. A basis of a vector space is a set of of linearly independent vectors that spend the entire Vector space. Every Vector in the space can be expressed as a unique linear combination of the basis vectors. So there are a couple of parts in this definition; they are really important and first thing that we need to uh mention here is this Vector space that says it is a set of linearly independent vectors that spend the entire Vector space. This is very important because um here we are with the basis is simply this Vector space that is a set of linearly independent vectors, which means that one of these vectors cannot be Rewritten as a linear combination of the other one, so we have a linearly

Independent vectors and they span the entire vector space. So, for instance, if we are in R<sup>2</sup>, then the basis of a vector space is a set of linearly independent vectors that span this entire R<sup>2</sup>. So, do we then need to have four vectors forming a basis? Let's say we have a basis of vector space. For us to say that this is the basis of this vector space, let's say in R<sup>2</sup>, we need to first prove that these vectors are linearly independent, and two, they span the entire R<sup>2</sup>, which means that the span of these vectors is equal to R<sup>2</sup>. We can actually be even more specific. In an example of, let's say, having vectors A and B, we can say that this set that we have here consisting of vectors A and B in R<sup>2</sup> form the bases of a vector space if the first criteria is that A and B are linearly independent, and the second criteria is that those two vectors together they span the entire vector space of R<sup>2</sup>, which means that the span of A and B vector space is equal to R<sup>2</sup>. On more specific example, and then the second part of this definition says that every vector in the space can be expressed as a unique linear combination of the basis vectors. Which means in our specific example when we had this A and B forming the bases of a vector space, this means that if we prove that this is indeed the basis of this vector space, then any combination, every vector, let's say a vector C that consists of these C<sub>1</sub> and C<sub>2</sub> elements, that this vector, this random vector from R<sup>2</sup>, C, can be represented as a linear combination of these vectors A and B. So, let's say we have a coefficient K<sub>1</sub> * A + K<sub>2</sub> * B. Then here we are representing this random vector C as a linear combination of these vectors A and B, which form the bases of a vector space of this vector space.

So we have previously spoken about the null space and column space. So let's now go ahead and do one more example when we are calculating the null space and the column space, and then we are again calculating this concept of basis of vector space and a basis of the column space, and then we will be uh finding the basis of a vector space uh with um R<sup>2</sup> example. So given that we have already looked into this concept, the basis of column space and basis of null space, I will try to uh go through this example a bit more quickly to save time on more complex concepts. So let's say we have an example of a matrix, and that matrix is A is equal to 1 2 3 6. This is our 2x2 matrix A, and the first thing that I want to do is to understand, look into my matrix and understand whether I'm dealing with unique vectors or not, and by unique I mean whether I'm dealing with two vectors that are linearly dependent or linearly independent. This kind of inspection always helps us to save time when we are doing our calculation for the null space and for the column space and for the basis of null space and basis of column space. Now here we can see that this is our A<sub>1</sub>, the first vector, the first column vector forming the matrix A, and this vector is the A<sub>2</sub>. Another thing that uh we can notice here is that we can easily take the first column A<sub>1</sub>, multiply it by two and get the A<sub>2</sub>, because 1 * 2 is 2, 3 * 2 is 6. That is that 2 * A<sub>1</sub> is equal to A<sub>2</sub>, which means that we can say that A<sub>1</sub> and A<sub>2</sub> are linearly dependent. Okay. So seeing this and knowing this, this can help us to quickly go through our calculations of the bases of the column space and the bases of a null space. So let's go ahead and first calculate what is the basis of null space of A. So we have already learned that the um basis of a null space can be calculated when looking into the first null space. So we see we need to calculate the null space and then we need to calculate the basis of that null space. So this means that we need to get the N<sub>A</sub>, and we have learned that in order to get N<sub>A</sub>, we for that need to solve the Ax = 0 problem, and this x will give us the null space of A. We have also learned that the null space of A is equal to the null space of RREF(A), which means that using Gaussian reduction or Gaussian elimination we can quickly find the solution to this problem of Ax = 0 and find this x. This is simply solving a similar problem, only in this case the B, so this is equal to zero because we are dealing with the homogeneous case. I won't do the calculation for this; we have done a ton of examples when we were doing this step-by-step calculation, getting the uh augmented matrix of A and then uh doing all these different row operations, normalizations and then eliminations in order to uh get this uh complex matrix A to the point of uh basic representation from which either we can visibly see the solution to the problem or we can at least simplify it and describe it as a linear combination of vectors. In this case, if you go ahead and solve this problem, you will find that the x that solves the Ax = 0 problem is unique, and this x is equal to -0.894 as the first element and then 0.447 as a second element. This can be a good practice also to refresh um the memory when it comes to the Gaussian elimination and reduction. The example itself is quite simple, the A is just a 2x2 matrix um and um by performing a couple of operations uh in terms of normalization and elimination you can find this X for your A = 0.

Given that now we know what is the solution to A = 0 problem, now we know what null space is because in this case this all helps us to understand that the null space of A is then equal to the set, the vector set where as part of this we got just a single column which is -0.894 and 0.447. This is the null space. This is the first part, I will say 1.1, and then 1.2 will be to get the basis of this N<sub>A</sub>, and we have just seen what is the definition of the bases. So the basis of vector space is a set of linearly independent vectors that span the entire vector space. Therefore, given that we got just this single vector as a solution to our problem, we can see then very quickly that the null space of A is based on this, and then the basis of the null space is simply this entire set. So knowing what the solution is to our homogeneous problem A = 0, so let me also write down in here then we know that the null space, the N<sub>A</sub> is then equal to the vector -0.894 and 0.447. This is my vector X that solves this Ax = 0 problem, and this is simply the null space of A, and given that we have calculated and we have got this unique solution to our problem, we can say that any vector in R<sup>2</sup> can be represented as a linear combination of this vector. So 1.2, any vector in R<sup>2</sup> can be represented as a linear combination of this vector X. Therefore, we are saying that the bases of null space of A is this entire set consisting of the single vector. So this is about the basis of a null space. Let's now quickly look into the concept of the basis of a column space. So the first thing we need to then uh get is the column space. So to get the basis of column space, we need to get the C<sub>A</sub> first, which is the column space of A, and what is the column space of A? The column space of A is the uh set, the space of the vectors that we can see in here in this A<sub>1</sub> and A<sub>2</sub>. It's quite straightforward, so these two vectors they form the column space of this matrix A. So then the C<sub>A</sub> is simply the set of {1 3} and then {2 6} vectors. This is A<sub>1</sub>, this is A<sub>2</sub>. Now we have just seen in the beginning before even starting our calculations that A<sub>1</sub> and A<sub>2</sub> are linearly dependent because A<sub>2</sub> can be written as 2 * A<sub>1</sub>. So one of these vectors can be written as a linear combination of the other one. This means that we got just a single linearly independent vector, and why is this important? Because we have seen in the definition of the basis that for us to have a basis we need to have linearly independent vectors. So the basis of vector space, in this case the column space, is a set of linearly independent vectors that need to span the entire vector space, in this case R<sup>2</sup>. So therefore we need to look into the C<sub>A</sub> that we got in here and select one of these two vectors that can be considered as linearly independent. Let's say we pick {1 3}. Now we know that we can then write any vector in R<sup>2</sup> as a linear combination of this vector {1 3}. So we can scale this vector {1 3} and get a new vector in R<sup>2</sup>. Therefore, or we are saying that the basis of column space, basis of column space of A is then the set of {1 3}, because {1 3}, so A<sub>1</sub> is then linearly independent and the span of A<sub>1</sub> is R<sup>2</sup>.

Now when it comes to the uh basis of the entire R<sup>2</sup>, one thing that we can notice is that this A<sub>1</sub>, so {1 3}, it's not forming, it's not spanning the entire R<sup>2</sup> because because we cannot uh write any random vector in R<sup>2</sup> as a linear combination of this two. Therefore, we are saying that this is the basis of column space, but we are not saying that this is the basis of R<sup>2</sup>. And the final element in this definition that I want you to uh focus on is that every vector in the space can be expressed as a unique linear combination of the basis vectors. So in here we have looked into this idea of bases of a null space and the basis of column space, and we saw that we are talking about specifically the null space and column space, but when it comes to the entire space, for instance the basis for R<sup>2</sup>, then the basis of column space, for instance, is no longer um helping us because the basis of column space it consists of this vector {1 3}, and this {1 3} alone is not satisfying the second criteria that says that this vector needs to span the entire vector space because this {1 3} vector it's a single vector and this vector it is not forming the entire R<sup>2</sup>, it's not um the basis for R<sup>2</sup>, it's not spanning the entire uh R<sup>2</sup>. So given that the {1 3} is not spanning the entire R<sup>2</sup>, because of that we know that the {1 3} is not the set of {1 3} is not the basis of R<sup>2</sup>. So this distinguishing of the basis of R<sup>2</sup>, basis of column space, basis of null space is really important because basis for R<sup>2</sup> it means that we need to find a set of linearly independent vectors that they together form the entire R<sup>2</sup>, they span the R<sup>2</sup>, which means any random vector that we can see in R<sup>2</sup> we can represent as a linear combination of the vectors in this space. So in here let me also prove that this {1 3} alone is actually not forming the R<sup>2</sup>, it's not spanning the R<sup>2</sup>, which then uh concludes that they are not the, it is not the basis of R<sup>2</sup>. And after this I will then provide you an example where we have a set of vectors that span R<sup>2</sup> and are linearly independent, which means that they are the bases of the entire R<sup>2</sup>. So first I want to show you why this single vector {1 3} is not the basis of R<sup>2</sup>. So being the basis of R<sup>2</sup>, we have the criteria that the vectors need to be linearly dependent. So let me actually clear up some space here. So I want to see and find the basis of R<sup>2</sup>. First I want to prove that this set, which is the basis of column space, I want to prove that this is not the basis of R<sup>2</sup>. Then I will also, as part of the second part of this proof, look into the case when we do have vectors and the set of vectors it forms the basis of R<sup>2</sup>. So the first thing, the first criteria of the basis of R<sup>2</sup> says that, quote 1.1, the first criteria says that this vector in this vector space, it needs to be, they need to be linearly independent. Well that criteria is valid given that {1 3} is linearly independent. This means that criteria one is satisfied. So whenever you got just one vector, this criteria is automatically satisfied. So then you have the 1.2 which says that we need to have this span of these vectors equal to R<sup>2</sup>. So is the span of {1 3} the R<sup>2</sup>? Well, no. And how we can prove that? Because the idea is that any vector, including an example where I have for instance uh let's say {4 5}, this vector that I need to be able to find a scalar that will help me to create a linear combination, let's say C, linear combination using this vector {1 3}, which will then set this amount, this to be equal to this. Which means that I need to be able to write my random vector {4 5} as a linear combination of this vector that forms my uh vector space. So let's see whether that is even possible. Well here I got 4 and 5. If I do this multiplication in the right hand side I get C and here I got 3C because C * 1 is C and 3 * C is 3C. And this means that I have an equation 4 = C and 5 = 3 * C. From this I get that the C = 4 and C = 5 / 3. But that is impossible because 4 is not equal to 5 / 3, which means that I'm proving in here and I got to prove that the uh any random chosen vector {4 5} cannot be written as a linear combination of this vector that forms this uh space. Therefore, a random vector from R<sup>2</sup> can't be written as a linear combination of {1 3}. Criteria two is not satisfied because for that we had to say that this span of {1 3} is equal to R<sup>2</sup>, which we saw that it's not the case.

Okay, so now we have proven that the {1 3} is not forming the basis of R<sup>2</sup>. Let's now look into what then does form the basis of R<sup>2</sup>, an example of it. So we are familiar with the unit vectors e<sub>1</sub> and e<sub>2</sub> in R<sup>2</sup> which form the identity matrix I, and this is {1 0} and this is {0 1}, also $\begin{bmatrix} 1 & 0 \\ 0 & 1 \end{bmatrix}$ in the form of a matrix. So in this example we have a set consisting of e<sub>1</sub> and e<sub>2</sub>, where this is this e<sub>1</sub>, this is the e<sub>2</sub>, and the set corresponding to this vector space is then {1 0} and then {0 1}. And now I will be proving that this space, this vector space does indeed equal to the bases of R<sup>2</sup>. So this is the basis of R<sup>2</sup>. So the first criteria of the bases is that these two vectors should be linearly independent. Now we can quickly uh remember from our previous theory that the two unit vectors {1 0} {0 1} are actually linearly independent. That's something that we have proven, and you can easily see it also from here. There is no way that you can find um scalar C that you can multiply this vector with and get a vector {0 1}, because for this one to become a zero you need to multiply this with zero, but then 0 * 0 is not equal to 1, which means that there is no way that you can find a scalar C to multiply this e<sub>1</sub> to get the e<sub>2</sub>. So let me write this down: e<sub>1</sub> and e<sub>2</sub> are linearly independent because there is no scalar C which is a real number such that such that C * e<sub>1</sub> = e<sub>2</sub>. So this means you can't write e<sub>2</sub> as a linear combination of e<sub>1</sub> or vice versa. This means that e<sub>1</sub> and e<sub>2</sub> are linearly independent and this satisfies our first criteria. So criteria one is satisfied. What we have also learned is that any vector in R<sup>2</sup> can be actually written as a linear combination of unit vectors that form that um R<sup>2</sup>, in this case {1 0} and {0 1}. So let's assume that this random vector is {C<sub>1</sub> C<sub>2</sub>}. So this is C vector, and what we want to prove is that we can always write this C in terms of linear combination of these two vectors. And how can we do that? Well, let's say here we got a K<sub>1</sub>, K<sub>1</sub> which is a real number, and we multiply this by {1 0}, and then we add K<sub>2</sub>, K<sub>2</sub>, and then here {0 1}. So this is our e<sub>1</sub>, this is our e<sub>2</sub>. Can we do this? Well, what is this? This is equal to {K<sub>1</sub> 0} + {0 K<sub>2</sub>}, and what does this give us? Well, this means this amount, let me write it over: K<sub>1</sub> * {1 0} which is the e<sub>1</sub> + K<sub>2</sub> * {0 1} which are which is our second vector e<sub>2</sub>. This is equal to {K<sub>1</sub> 0} + {0 K<sub>2</sub>}, and this is equal to {K<sub>1</sub> K<sub>2</sub>}. So I got on one hand this vector {C<sub>1</sub> C<sub>2</sub>} which I want to write as a linear combination of K<sub>1</sub>e<sub>1</sub> + K<sub>2</sub>e<sub>2</sub>. If I take the K<sub>1</sub> equal to C<sub>1</sub> and K<sub>2</sub> equal to C<sub>2</sub>, well then in that case I can prove, so this is basically equal to {C<sub>1</sub> C<sub>2</sub>}, which means if I take this K<sub>1</sub> and K<sub>2</sub> equal to C<sub>1</sub> and C<sub>2</sub> respectively, and those numbers are given, then I can represent this vector C as a linear combination of e<sub>1</sub> and e<sub>2</sub>, which is what I had to prove in order to say that the span of {1 0}, which is the e<sub>1</sub>, and {0 1}, which is e<sub>2</sub>, is equal to R<sup>2</sup>, because any random vector that will be provided to me with an element C<sub>1</sub> and C<sub>2</sub>, and those are just real numbers, can be written as a linear combination of these two vectors. This means that the span of these two vectors is equal to R<sup>2</sup>, and this is basically the second criteria. So criteria two satisfied, and if the criteria one and criteria two are both satisfied, it means that this vector space of {1 0} and {0 1}, this is the basis of the entire R<sup>2</sup>.

So let's now talk about the concept of projections. By definition, a projection of a vector A onto another vector B is the orthogonal projection of A along B. It's denoted by proj<sub>B</sub>A. So projection of A onto B. So here is the A and here is the B, and represents the component of A in the direction of B. So component of A in the direction of B. All right. So in order to properly understand this concept, the intuition of it, let's actually make use of the R<sup>2</sup> space. So let's first start by picturing in our flat world the R<sup>2</sup> coordinate, so the Cartesian coordinate system. So let's say here we got our y-axis, here we got our x-axis. So this is the X, this is the Y, and uh here we of course we need to keep in mind this is just an example. When it comes to projections we can always go beyond R<sup>2</sup>, but for keep it simple and truly understand this concepts and this intuition behind the projection I want to simplify this and do the example in R<sup>2</sup>. So here uh imagine that we got this line, and this is our A line that goes through the center, that let's call this line B. So B is a line in R<sup>2</sup>. Let's say this is that line, and now that imagine that we have this vector which is part of this line. Let's say this is this line, and this line is representing by uh on this line we got this vector B, and this vector is basically part of that line, as you can see. This is the vector B on this line B. So we know from this concept of the line spanning the R<sup>2</sup> and then vectors, we know that in this case, independent what is the magnitude of this vector, what is the direction of this vector, we can represent this line B by this linear combination based on this vector. So linear combination of this vector, which is in this case the set, then here we got some C, where C is a real number, multiplied by this vector B, knowing that this C is just a real number. So let's make it actually green. So

We can basically say that this entire line B can be represented as this set of linear combinations of these vectors. So, for instance, if this is one and we do the C is equal to two, then we can get this part of, so we can get this vector. Otherwise, this is equal to three, we can get this vector, or C is equal to four, this vector, and then and so on. Which means that we can always come up with a linear combination forming a part of this line. Therefore, we are seeing that this line can be represented as all these linear combinations of this vector B, which is part of this line, and here the C is just a scalar, so a number which is a real number. So this C * vector B represents this entire line. We will knit this in a bit, but for now, imagine this line and part of this line which is this vector B.

So imagine then that we got yet another vector, which is let's say in here again going from the center, but this time in this different direction, so in here. This is vector a; we call this vector an A. So you can see that this vector a is actually much longer than the vector B, and we see that vector a is not lying on the same line as B. So B is lying on the line B, and a is not lying on the line B. Now let's say we want to project this vector a onto this vector B, which means that we want to project this a in this direction. So we want to bring this vector a onto this line. Let me actually use a different color, and the word of the projection actually does make sense in here as you might notice because we're trying to cast the shadow of a onto this line of B.

And how can we do that? We can only do that if we connect this vector a like this with this orthogonal line, let's say by using a different color of this, so with this perpendicular line we then will be connecting the vector a to the line B, because we want to project our vector a onto this direction. So this perpendicular line that you see in here that goes from vector a to the line B, where on line B we have the vector B. So here is the line a, line B, and this perpendicular line, it goes from a to line B, and on line B we have the vector B that is represented like this. Then the projection of a onto line B is this shadow vector that you see in here, and the word projection or the name projection actually does make sense because we are projecting this vector a onto this line and it creates this shadow. So we are casting this shadow on here, and this vector is what we are referring to as projection of vector a onto line B. Notice that we don't say projection of B on vector B, but instead we are saying projection of a on the line B.

Then another thing we can notice is that we are getting this projection of a on B, so this vector, by taking the vector a, so vector a, and subtracting from that projection of a on B. That is the formula for this vector that we refer to as a perpendicular that goes from a to line B. So when drawing this perpendicular line from a to line B, we are referring to this as a minus projection of a, b, because you can see that this vector is simply this vector minus this vector. That is the mathematical expression for this perpendicular line. So how we can then find out what is this C that we got in here? Because we understand that to get this exact formula for the projection of a on the line B, we need to understand what is the scalar, specifically what value are we using to multiply this vector B to get to this point. So what is that c? What is C? What is C such that c times a is then equal to projection of a on the line B? Because we can have different sorts of a linear combination of vector B on this line B, and in fact, B, this line B is the set of all linear combinations of this vector B, and I want to know specifically what is the vector that we see in here, what is the shadow vector, because this is the projection of a on the line B, what we see in here.

Now how can we do that? Well, let's first formally define on this specific case what is the projection of a on this line B. So projection of a on line B is some vector that is also on line B, where a minus projection of a on B is perpendicular or orthogonal to this. This is basically the definition of the projection of a on line B under this specific example. So in this case, the way we can find this projection is by looking into this C. So this is what we are interested in, this specific, specific C * B vector, and knowing C and knowing B, we already know what is B, what B is; knowing C, we can then describe this specific projection. So one thing that we can know is the condition under which we say two vectors are orthogonal; that's something that we already have learned as part of the previous lessons. So let's go ahead and find that amount. So now what we need to do is to calculate this value of C, because value of C calculation will then lead us to the exact vector that we are interested in, which is this projection. So our end goal is to find out what is this projection of a on B; this is what we want, and for that we need to calculate this C because we already know the vector B. So let me quickly remove this part, cuz here we will then do our calculation.

So one thing that we need to make use of is this part when it says orthogonal, because we know that if two vectors are orthogonal, then their dot product is equal to zero. So we know that this vector is orthogonal to this target vector, which means that we can say that the vector a and then minus projection of a on B multiplied with vector B, that this is equal to zero. This is something that we know by definition of orthogonality: two vectors are orthogonal, it means that their dot product is then equal to zero. Now let's make use of that part. So this means that we need to describe this projection of A and B; we need to make use of the fact that we know that this projection of a onto B is actually some linear combination of vector B. So let me actually go ahead and remove this part; we already know the definition. So let us go ahead and calculate that C that we need in order to find out what is this entire projection. So few things that we need to clear out are those formulas, because then we can make use of them to find the C. So we know that by definition the projection of a on the line B, it is this vector that we get where we draw this perpendicular line from vector a onto line B, and we said that this line is equal to this amount. This is simply the vector a minus this vector, the shadow vector, which we said it's defined by projection of A and B, this thing. So we can make use of that because we also see in here that this we are saying is orthogonal to this vector. So given that this vector a minus projection a, b is orthogonal to line B, that is also orthogonal on this specific vector, which is the projection itself. So from this we can make use of the fact that two vectors, when they are orthogonal, their dot product is equal to zero in order to find this value of C.

So firstly we just set that the a minus projection of a on the line B, that this multiplied by this vector B is equal to zero, because those two lines they should be perpendicular. But at the same time we know that this is simply the linear combination of this vector, because this line is perpendicular to this one, and this line is some linear combination of this vector B. Because if I have here a vector and then I have the longer version of that vector on the same line, which is then a linear combination of this original vector, let's say this is my vector B, then this second vector that I have in here is then equal to some C * vector B. This is also exactly what we said in here. Here we said any vector on line B can be represented as a linear combination of vector B, and this is exactly what we are seeing in here. So this projection is simply that C times vector B; this is something that we have already said. So we are just making use of that to fill in that volume. So this then results in a minus this C * B multiplied by this vector B is equal to zero formula. So here we are simply making use of the fact that the projection of a onto B is the shadow vector, which is then equal to some linear combination of this original vector B, which is on this line B. Then I can easily find the scalar C from here because we know how we can easily calculate this dot product. So let us actually go ahead and do that. Let's first multiply this a by B and then minus, so I'm simply opening the parenthesis C * then I got B by B and this equal to zero. So C * B time B is then equal to a and b, which means that c is equal to a * B / to B * B.

Now when we have the C, we can easily derive the formula for the projection of a on the line B. So this is the first part, this is the second part. So then the projection of a on B, so projection of a on B is equal to this c, c times the B, and we just found out that this is equal to the C was equal to a * B / to B * B, and now we need to take this C and then multiply by vector B. This is then the projection of a on B, this vector. So projection of a on line B. So you will notice that this is the same that we just got. So whether you compute the projection of a on the entire line b or projection of a on the specific vector b, as we are using the vector b as a source for drawing our line, this is the same as the projection of vector a on vector B, and this is the same formula as we see in here. So this is the projection formula that we have just found out. So projection of a onto B is given by this formula: a * B, so the dot product of the vector A and B divided to the dot product of the B with itself and multiply with the vector B, and this is the in here. This is something that we have calculated time and time again in our examples.

So if we go back to our example, then here we can see that this is our vector B, this is our vector a, and we are saying if we take the vector a and we project it onto this vector B, then we can calculate this projection, which is in here the formula for this entire vector, which we are calling projection of a on b or projection of a on B. This can be found out. So the length of that vector we can find by using this formula: so the dot product of vector A and B divided to the dot product of B with itself and then multiplied with vector B. So again, a dot product, and this is of course something that we get as a vector. So this is a vector, something that is equal to this entire vector in here, this vector. So I know that this might look a bit messy because it contains many moving parts, but I wanted to provide this detailed explanation and the step-by-step process even if it is a bit confusing and a bit messy in the beginning, because this helps us to understand what this formula is about and what is the intuition behind it. Because what we are doing is that we are making use of the fact that the line can be represented as a linear combination of all the vectors that we use in here. So this is vector B, and this entire line B is a linear combination of this vector B, and we can make use of that in order to find that scalar that we are multiplying to create this single linear combination that will end up giving us this vector that we see in here, which is the projection, the projection that we are interested in, which is this line. This is the projection that we are defining by this projection a onto B, and we can get that by making use of the fact that this, this perpendicular line that we are creating in here, which is simply the vector a minus this projection, this is this vector, this projection vector, that this is perpendicular to this line B, and if the vector B is part of this line B, this means also that this line a minus projection a is also perpendicular to that vector, vector B. Making use of that formula, we can then make use of the product of the two; we know that the dot product of two perpendicular vectors is equal to zero. Making use of that, we can then obtain this specific scalar C that we need in order to get the final formula for our projection. We are interested in this C because knowing C we can then multiply with this vector B to get our final projection, and we have found that that projection a on B is defined as the dot product of the A and B divided to the dot product of the B with the B and multiply with the vector B, and this is again a vector.

Now let's look into a couple of numeric examples to clarify this topic and practice with it. So given vectors A and vectors B, find the projection of a onto B. So without looking into the answer, I will quickly go onto that example itself. So vector a is this vector 3, 4; can also represent this by our more common notation, which is three and four, and then vector B is one and zero. So let's quickly draw our coordinate system; this our x-axis, this our y-axis, and then what is the a? The a is three and four, three and four. So this is our a, and what is the B? The B is one and zero, which means that our line B is then C times the vector B, given that the C is a real number. And one thing that you can notice is that the line B is actually our x-axis; it is this line. This is our line B; this is our line l, b. So the projection is then this line; this is our projection because we can know that by drawing a perpendicular line in here from a to the line B, we can get then the connection between our vector a and vector B and create our projection. So this is then the a minus projection of a on line B, and this part is then, this is then this projection a on B. And how we can get this projection? Well, we just learned that the projection of a on B is equal to dot product of a with B divided to dot product of B with B itself and multiply it by B. This is the formula that we can use, and even if you don't remember the formula by heart, you can make use of this visualization to figure out what that formula is, because we know that if this line is perpendicular to this one, then a minus projection of a on B multiplied by this projection a on B should be equal to zero, and this projection of a on B is equal to some scalar C multiplied by vector B; that's something that we see in here.

The first thing we need to do to compute the dot product between a and b, a * B is equal to 3, 4 multiplied by 1, 0. This is the dot product, which is then equal to 3 * 1 + 0 * 4, and this is equal to 3. The next thing we need to do is to compute the dot product between B itself, so B * B, and what's that? That is 1, 0 with 1, 0 multiplied; this is equal to 1 * 1 + 0 + 0 * 0 is equal to 1. Then the third thing that we can do then is to obtain the final value, which is projection of a on B is then equal to three divided to 1, multiplied by the vector B which is 1, 0, which is equal to 3, 0. And this actually makes sense visually too, as you can see in here. This is the three for the x-axis, and here we have the center 0. So this projection is then the vector 3, 0. So even without calculation, we could see just from plotting the on the coordinate system the vectors A and B that the projection of a on B will be this vector 3, but we have followed the formula in order to do calculation step by step, which is something that you can see in this answer too. So the projection of this vector a onto B is then this vector of a length three in the direction of B. So you can see that it is of the length of three. So this is the three on the direction of B, so on the line B.

Let's now move ahead and look into a different example, but this time we will do the calculation in a quicker way. So we got two vectors, 4, 3, and B is equal to 2, 0, and we need to find this projection of a onto B. So the first thing we need to do is to calculate the a * B, which is equal to 4, 3 multiplied by 2, 0, and that's equal to 4 * 2 + 3 * 0, and it's equal to 8. The second thing we need to calculate is the B dot product with B, which is equal to 2, 0, 2, 0; this is then equal to four. And the final part is to take and from this one and two this values and then bring them all together. So then the projection of a on B is equal to 8 divided to 4 multiplied by the vector 2, 0, and this is equal to 8 / 4 is 2, 2 * 2 is 4, 2 * 0 is 0. So we are getting this 4, 0 vector. So projection of a on B is then this vector 4, 0, which is actually on this x-axis, similar to what we had before, only with the length of toward the direction of, which is then equal to 4 and 0. And this is again similar to what we had before where we got the projection of a on B on that end up on the x-axis, but now with the length of four. So now our projection has the following vector, so the following magnitude and direction. So this is the step-by-step process that I just followed if you want to do it a bit slowly, and this is the final result. So the interpretation of this projection is that this projection a onto B is simply this 4, 0. This means that the A's component in the direction of B it spends four units along this x-axis that we saw in here, because this is the value X, this is the value of y. So this projection shows us that A's influence in the direction of B is completely horizontal with this magnitude of four, because we saw that we end up with the projection on the x-axis again. So this was four, this was our projection vector, and if you plot this entire vector a and vector B on this x-axis and y-axis, then you can clearly see that the horizontal line that we end up with the projection of a and b is very similar to what we had before in here.

Let's now talk about a concept of orthonormal bases. So let's now define what the orthonormal bases are. So by definition, orthonormal basis for a vector space is a basis where all vector vors are orthogonal or perpendicular to each other, and each vector is of unit length. So as you can notice here, here we have a special type of basis; it's called orthonormal basis, because in the beginning of this section of this module we defined formally this concept of bases; we talked about the concept of colal space and then the basis of a comp space, the null space, the basis of an null space, and then we talked about the concept of the bases of the entire space, for instance the R2, and now we are defining a special type of bases which we are...

Referring to orthonormal bases, and this orthonormal basis as you can see from this definition, it contains two criteria for it to be orthonormal. So an orthonormal basis for a vector space is a basis where: A) all vectors are orthogonal or perpendicular to each other; and B) each vector is of unit length. We already have learned that when we have vectors, let's say Vector A and B perpendicular, it means that A and B, their dot product is equal to zero. That's the first criteria that we need for calling our basis an orthonormal basis. Then the second criteria is that each of these vectors they need to have a length of one. If we have this condition satisfied, then we are saying that our vectors they help us to form this orthonormal basis.

If we got three vectors forming this vector space, it means that we need to have A * B = 0, A * C = 0, and then B * C = 0. This is if we are in, in case we are using three different vectors that define our vector space. In this case, let me make this part smaller, so let's put the length of B in here. In this case, the second criteria becomes that the length of A is equal to the length of B and then is equal to the length of C and is equal to one. So depending on the number of vectors that you use to form your vector space, the proof that you are dealing with orthonormal bases will be different. Here we got just two vectors; here we got three vectors. But in both cases, we first need to prove that we are dealing with uh vectors, each of which are a set of orthogonal, perpendicular vectors, and all of them pairwise they need to be perpendicular. And at the same time, the second criteria says that they all need to have a unit length, so their length should be equal to one.

We need this orthonormal basis in order to simplify our calculations, including the calculations of projections and transformations that we just saw before when we were discussing this concept of projecting a vector onto a line or projecting a vector onto not a vector, because we were in this basic case when we had just two vectors in R2. And calculating projection in R2 is very easy because we can make use of this formula: A and then B, the dot product of them, and then divided by the product of the B and then times the B. This was quite straightforward, right? But when it came, so this is the projection of A on B, but when it comes to projection in higher dimensional space, let's say you have R5 or you have R100 or R1000, then it becomes much more difficult to do those projections and to calculate the projections. And for those cases, we can make use of this concept of orthonormal basis to simplify our calculations, and we will see that in a bit. So let's first understand this orthogonality and the normalization part. Orthogonality refers then to the part of uh when we are saying that the vectors should be orthogonal to each other, and the normalization refers to the fact, uh to the fact that the length should be one. This is basically the set of two criteria that I just discussed. This is the summary slide that will give you an indication what is meant by that. So if we have two vectors V and W, then we say that the first criteria is that those two vectors are orthogonal, which means their dot product is equal to zero, and we are saying that their length is equal to one, which we are referring as a normalized vector. So if the vector has a length of one, then we are calling a vector V normalized. So if both of this criteria of normalization and orthogonality is satisfied, that we are saying that we are dealing with an orthonormal basis.

So now where we have learned this idea of projections, also this idea of orthogonalization and the uh concept of orthonormal basis, we are ready to discuss the concept of the Gram-Schmidt process. The Gram-Schmidt process is this method for orthogonalizing a set of vectors in an inner product space and turning them into an orthonormal set. So let's say we have a set of vectors, we want to uh bring and transform all these vectors onto this orthonormal set of vectors, which means that we want them to be orthogonalized, so we want them to be perpendicular, and we want them to be normalized, because we know that the two criteria were specified, right? So the first criteria was that we need to have vectors orthogonal; hence we are doing orthogonalization. And the second criteria was that they need to be normalized because we want the vectors to have length one, so we are doing normalization. This process of turning this set of vectors onto this orthonormal set by using this method of orthogonalization, which is something that we are referring as a Gram-Schmidt process. This is something that we can use in order to simplify later these different sorts of transformations which we need in order to perform bit more advanced uh transformations like matrix uh factorization, different decomposition techniques. Given this set of linearly independent vectors, this process which we are referring as Gram-Schmidt process produces this orthonormal set that is spanning the same subspace. So we have the same subspace; it's just that we are turning the set of vectors into an orthonormal set of vectors that is spanning the same subspace.

The Gram-Schmidt process step by step looks like something like this. So given the vectors A1, A2 up to An, the first thing we need to do is to start with the vector V1 which is equal to our first vector A1, and first we need to normalize this vector. And how we can normalize this vector? Well, we need to take this vector and we need to divide it to its length. So the Gram-Schmidt process step by step will look like, like something like this. So in the first step, what we need to do when starting with these vectors of A1, A2 up to An, so in Rn, we need to first take the first vector and we need to normalize it. And how we can normalize the vector and ensure that its length is equal to this length of V1? Well, we need to take that vector and we need to divide it to this length, because when we take the vector V1 and we divide it to its length of V1, then we will ensure that the length of that vector is equal to one. We can actually prove that very easily, but I won't do it in here. Uh feel free to go through the process, assuming that the length of the vector, what, what you want to achieve at the end is that the length of a vector V is equal to one. This is something that we want to achieve, and this normalization process can be done if we find a way to ensure that we uh get this E1, because E1 means that we end up with this vector 1 0 0 0 0 0 0. This will be for first vector, so V1, this is E1. So the one is really important here. So we want to normalize this vector V1 by uh ensuring that we get the E1, so we go from V1 to E1, and the way we do that is that we take the V1 and we divide it to the length of V1, and in this way we get the E1. So the normalized version of V1 is E1.

So then for each subsequent vector Ak, which means A2, A3, A4 up to An, we need to subtract its projection on all the previously computed orthogonal vectors. In this way, by using this step two, we are ensuring that all these different, each pairwise set of A1, A2 and then A2, A3 etc., they are all orthogonal to each other. And we know that this projection is something that we got when we had this two perpendicular lines. So we had this vector, we're projecting onto this vector, we got that by finding this perpendicular line and making use of that, using this property, we are then making use of that in order to see how we can ensure that the subsequent vector that we have is always perpendicular to this one. So let me actually write down what is in this formula. So here Vk is equal to Ak minus the sum of all the projections. So then we need to normalize the Vk to get the Ek, and then we need to repeat this step two and three for all vectors, which means that first here we apply this normalization on the vector A1, so V1 is to A1, and then we get the normalization by getting this E1. So E1 is normalized version, and then we need to apply a bit different tactic for our V2, V3 up to Vn, and then let me actually write down this for this general case. So what this process, this the Gram, let me ensure that I'm not making a typo, Schmidt process step by step means: step number one, for vectors A1, A2, A3…An, so we are in the Rn, then step number one basically says: take the V1 and set it equal to this first element V1, this is A1. Then what we need to do is to normalize it to get the E1, so normalize V1 to get E1, which is equal to 1 0 0 0 and then…zero, and the size of this n by one. And how we can do that by taking this vector V1 and divided it to the length of V1, which basically means in this specific case A1 divided to the length of A1. This will then give us our A1, this vector, this is basically what the step one entails.

Then in the step number two, we have for each subsequent Ak, where K is just an index referring to whether we are dealing with K is equal to 2, so A2, A3 and then…An, this is what basically the K is used for, to refer to which vector we are dealing with. We need to subtract its projection on all previously computed orthogonal vectors by using this formula. So let's actually do a couple of those cases to see what is going on. For instance, for K is equal to 2, so K is equal to 2, and here is the formula by the way, so Vk, Vk is equal to Ak minus some K is equal to, K starts with one, and then let me use a different index, so i is = 1 till K - 1, and then projection of Ak, Ak on Ei or Eii. So the Ei that we have just computed, because every time you are then normalizing and normalizing every time your vectors, and then you are uh finding out what is the projection of your vector A onto that Ei, and then you are subtracting that from your vector. So what this means in practical terms, when for instance your K is equal to 2, it means that V2 is equal to A2 minus sum of i is = 1 and then K is = 2, K - 1, this means this is one projection of A, A and then 2 on A1, given that this is one, this is simply equal to A2 minus projection of A1, that's normalized version of E1, and then A2, so projection of A2 on E1. And then in the step number three, we need to do, we need to go from Vk to get Ek. So basically we are ensuring with the step number two the orthogonality, uh orthogonality condition, and with step number three, I me add some space in here. So in the step number three, step number three, we then saying: let's normalize, normalize this Vk that we have just computed in here, because we remember that the second criteria after tonality is normalization that the unit or the length of the vector should be equal to one. So then Vk, in this case for K is equal to 2, for K is equal to 2 means that we need to go from V2 to E2, and the way we can do that is by taking the V2 by V2 and then divide it to the length of V2. This will then give us the E2. This is the normalization part. And the step number four basically means: repeat, repeat step two and three for all Ks, which means that if we go back, so we are done with V2, so we have obtained V2 and then we have obtained normalized version of V2 by getting this E2. We are ready to come back and do the same for K is equal to three, and for K is equal to three in step number two, we got V3 is equal to A3 minus, making use of this formula, sum overall i is = 1, K - 1 is = 2, and then projection of this time A three, three, K is equal to three and then on E2, actually it says Ei, let me remove this, this otherwise we would have made a mistake, this should be i because an i will change per K. This is the entire idea, we need to re- uh subtract all the uh projections. What this basically means is that we need to take A3 and this time given that here we have two instead of one in here, we need, you have an extra step, which means A3 minus and then what this formula basically says, this is the sum of the projections of A3 on Ei where i goes from one till two, so projection of A3 on A1 when K, so when i, this is the i is equal to one case, plus projection of A3 on A2, this is the i equal to 2 case. This is basically what this summation says, this is this element, and we have seen this as part of the high school but also the pre-algebra course. Okay, so now when we are clear on how we can calculate the V3 in the step number two for K is equal to 3, we are ready to go onto the step number three, and what was step number three? The step number three for K is equal to 3 was saying: let's take the V3 and normalize it to go from V3 to E3, and how we can do that by taking the V3 and dividing it to the length of V3 to get on to E3, and this cycle goes on and on until we cover all the Ks, so all the vectors. So the idea is that we first for our initial step we set the V1 equal to A1, we normalize it, then starting from the K is equal to two, we don't go first on and on, we orthogonalize it by formula in here, by using this we can ensure that each of these vectors is then orthogonal to all the other vectors. So for K is equal to 2, we ensure that this uh vector that we get is orthogonal to all the other ones, and the K is equal to three, that the third vector is orthogonal to all the other ones, and we are doing that in step number two. So for each K, for each K, we basically are ensuring that in this case we have an A vector that is orthogonal to all the other vectors in this set, and for each vector we are also normalizing it to satisfy the second criteria, because we had these two criterias to create this orthonormal set. So we are doing this in subsequent uh way, so first for K is equal to 1, so basically for A1, and then we are doing this for K is equal to 2, so A2, and then until K is equal to n, so An. What we are doing every time is that we are obtaining this V1 and then we go from V1 to E1 to normalize it, and then here we are getting the V2 here to go to E2 by normalizing it, so this basically the step two and step three, and then we do this every time up until to the point of obtaining Vn and then from Vn we go to En to normalize it. So this is the idea of this entire process step by step: to start with V1 as part of the step number one, and then as part of step number two for each subsequent vector Ak, so K is equal to two, obtain the Vk and then normalize it, for K is equal to 3, obtain the V3 and then normalize it to get E3, up to the point of the last vector which is An, the vector An, we compute the Vn and then we normalize it to get the En. And this is what this part is, which is the step number two that says: repeat steps 2 and 3 for all vectors. It means that every time when you increase your K, when you go to the next vector, we first compute the V, so Vk, and then you normalize it, you get the Ek, and then you go back to the step number two and three because you then again need to calculate the Vk and then Ek and then for the next K. So this is something that you will see also a lot when you are writing the code for your uh algorithms, because in many cases you need to do this repetition of the steps, so uh you for one vector you do something or for one iteration you do process and then you uh go back and do for the next one and for the next one. This process is what we are referring by: repeat step number two and three for all vectors.

So let's now look into an example. Let's apply this Gram-Schmidt process to vectors A1 and A2 where A1 is 1 1 0 and A2 is 1 0 1. So let's go ahead and do that. So A1 is equal to 1 1 0, A2 is equal to 1 0 1. We want to apply this Gram-Schmidt process to create this orthonormal basis for the subspace that is spanned by A1 and A2. So now we have this set 1 1 0 and 1 1 and what we want is to create an orthonormal basis. Creating, creating orthonormal basis which Gram-Schmidt process. So here we got only two vectors, so obviously it's this and it's a very simplified version of it. What was the first step in our case, uh in our algorithm? It was to set the V1 equal to A1. What we need to do, step number one, we need to set the V1 equal to A1 and we need to normalize, normalize the V1 to get E1, that's what our goal is. So let's go ahead and do that. V1 is equal to A1 and is equal to 1 1 0. That's our A1, so 1 1 0, and in order to normalize V1 and get the E1, we know that this is equal to V1 divided to the length of V1, which is then equal to take the V1, so that is 1 1 0 and then divide it to the length of V1, and you can very quickly see that given V1 is equal to V1 * V1, that's something that we learned in the very beginning of our fundamentals to linear algebra course, that the length of V1 is simply the dot product between uh V1 and V1, and it's equal to 1 1 0 * 1 1 0, which is equal to 1 + 1, so 2. So this is then equal to 1 1 0 / 2, which is equal to 1/2, 1/2, and then 0. This is our E1. So we are done with our step number one because now we have V1 and we got E1. So what was the step number two? In the step number two, we need to set the K equal to 2, this is the next K. So for A2, what we need to do is we want to get V2 and normalize V2 by getting E2, and how can we get that? Well, first let's find what is the V2. Well, V2 was, and using that formula that we saw before, which was this formula, so it's equal to Ak minus and the sum i is equal to 1, K minus 1, and then projection of A onto Ei, i. So let's take this formula over, this is equal to A2 because K is equal to 2, A, Ak minus su, and then i is equal to one till K - 1, and then K - 1 which is equal to basically 1 given that K is equal to 2, and then projection of A2 onto Ei, and this is equal to A2 minus, given that we got K - 1 is equal to 1, so the limit for our summation is equal to 1, so this one, this means that like before we got just one part as part of our summation, so minus and then projection of, let me actually keep the same color, I want it to be consistent, so projection of A2 on the E1. So you see here the i, i is equal to 1 and the limit of the i is K - 1 which is equal to 1, so we got here just E1. So we got the V2 formula, we can then now calculate it because we know…

A2 and the A2 is this: so one one 1 0 1. But now we got a problem; we don't know what this is, so let's quickly go and calculate this part. So projection of A2 on a1, and we learned from the projection formula that this is equal to H2 * E1 / 2 E1 * E1. So the dot product multiplied by E1, and what is this? This is equal to 1 1 multiplied by, and what is the E1? E1 we just calculated in here, so it is 1/√2, 1/√2, and then zero here / two, and then 1/2, 1/2, and then 0 0 multiplied by 1/2, 1/√2, and then Z here multiplied by the same vector, so E1. So this two cancel out; this two also cancel out. And as you can see, we are getting that the projection of A2 on E1 is equal to this vector. We can also manually check that actually. So let's let's do that. So let's see, we are not canceling out these vectors and instead we are manually calculating this. So here we got 1 0 1 multiplied by 0.5 and 0.5 0. This is equal to 1 * 1/2 is 1/2; 0 * 1/2 is 0; 1 * 0 is 0. So + 1/2. This amount is 1/4 + 1/4. This multiplied by the vector 1/2, 1/2, and then 0 in here. This is equal to 1/2; 1 + 1/2 is = to 3/2, and then 1; 1/4 + 1/4 is equal to 1/2. 2 multiplied by 1 2 and then 1 2 and then zero. What is this amount? Well, those two cancel out, so we end up with three * and then 1/2 2 and then 1/2 and then zero. This is then the projection 3/2, 3/2, and then zero. So let me remove all these calculations, and then we can take over the projection value which is 3/2, 3/2, and then zero to get our vector V2, which is equal to 1 - 3/2, 0 - 3/2, and then 1 - 0. And this is equal to here it is 1; here it is -3/2; and here is minus and then 1/2 2 because 3/2 is minus uh it is 1.5, and then 1 - 1.5 is simply -0.5. So this is then the vector V2. Then what we need to do is to normalize this vector to get DE2, which is then equal to V2 divided by V2 length, which is simply equal to V2 / √V2 * V2, so the dot product, and this is equal to let's take the V2 which is -1/2 and then -3/2 and then 1 and then divided by 2 and this amount. Let's quickly calculate that it is equal to so the length of V2 is equal to -1/2^2 + -3/2^2 + 1. This is equal to 1/4 + 9/4 + 1, which is 4/4, and then this is equal to 1 + 9 is 10; 10 + 4 is 14, so 14/4. This is the length of it, so 14/4. So then this is equal to this vector to this three-dimensional vector min-1 * √14 - 1/2 * √14/4 is equal to this is 7, so -7/√14 and then -3/2. Think I just made a mistake here actually. So min -1/2, so the first element and then divided by √14/4 is actually actually equal to this multiplied by 4 divided by 14. So you take this element then divide it to this one, and we know that a/b * c/d is equal to a * d and then b * c. So we are basically flipping this side; this is from pre-algebra. And then here this is equal to 2 and then is equal to -1/√7. Then let's do the second one, two. So we got -3/2 / √14/4 is actually equal to -3/2 * 4/√14, and then if we remove this this is then 2 this is 7; this cancels out this equal to -3/√7, -3/√7, and then finally we got 1/√14/4, which is equal to 4/√14; this equal to 2/√7, so 2/√7, and this is our A2. And given that we got just two vectors, so we have already reached the end of our solution. So now when we have already the V1 and the V2, the E1 and the E2, we have basically completed the process of this Gram-Schmidt procedure because we have already only two vectors, that means that we need to have V1 and V2 and then E1 and E2, and this is all that you need in case you got two vectors. If you have three vectors, of course, the process will include the same process of Step number two and three, so the V2 and the normalization of it two times for your k is equal to 2 and k is equal to 3. And then if you have more vectors, than every time you will have more of the steps, but at the end what we want to have is the set of vectors that are orthogonal and at the same time they are normalized. In this case, we say that this vectors form this orthonormal basis. Now why is this important? The applications of orthonormal bases: well, firstly, it simplifies complex vector operations, and this is the basis of many more difficult mathematical concepts like Fourier series or quantum mechanics. It's used also when when it comes to this orthonormal basis also signal processing, and it's a critical process in numerical methods, especially in machine learning algorithms and in data compression. So we will see this process to be used also as part of decomposition techniques, which is really important when it comes to different algorithms, whether it's optimization algorithms but also algorithms that are used for recommender systems, for example. And those concepts they all come together, and we will see later on when we will be discussing the concepts of decompositions and matrix factorization. So this orthonormal basis and this Gram-Schmidt process, they are really foundational in linear algebra; they provide tools for simplifying and also solving these higher-dimensional problems efficiently. Their application includes different fields of science and engineering, demonstrating their versatility and utility.

Let's now talk about the special matrices and their properties. So we are going to talk about special matrices like symmetric matrices and their example, diagonal matrices and their corresponding example, but also the orthogonal matrices with the corresponding example. So when it comes to the special matrices, special matrices have unique properties such as being symmetric or all nonzero elements on the diagonal like diagonal matrices or orthogonality in matrices, which means that we have orthogonal matrices. So when it comes to the symmetric matrix, it means that the matrix A is equal to its transpose, the A. So A is equal to A. In this case, we can confirm and say that the matrix A is symmetric. So in this case we have matrix A, and we know that the way we need to to transpose this matrix is to taking this rows and making them the columns of our transpose matrix. So A is then equal to 2 -1 and then zero; then the second row which is -1 and then 2 and then -1; and then the third row which is 0 -1 and 2. So the third row then becomes my third column. So as you can see, those two are the same. So I'm using then the definition of the transpose of the matrix, and then here then we end up with two matrices; they are actually the same. So we can see that the A and the A, in both the first column they got 2 -1 and then 0; the second column -1 2 and -1; the third column 0 -1 and 2. So their columns and their rows they are the same, which means that we are dealing with a symmetric matrix. So whenever we want to check whether the matrix is symmetric, we just need to take the transpose of it and see where the the matrix is equal to its transpose; in that case we are dealing with symmetric matrix. Do also note, therefore, for a matrix to be symmetric, it needs to be a square matrix. So it needs to be 2x2 in two-dimensional space or 3x3 in the three-dimensional space or nxn in n-dimensional space, which means that the number of rows should be equal to number of columns, because otherwise when you flip your number of rows with number of columns on in case there is no square version of that matrix, so m is not equal to n, in that case A will have a dimension of mxn, and then A transpose, so A transpose will have a dimension of nxm, which means that there is no way that A can be equal to A. This is not then possible; therefore, we need to have a square matrix for them to be symmetric.

Let's now talk about diagonal matrix. So a diagonal matrix has a nonzero element only on its diagonal, which means that in this case we have this nonzero elements on the diagonal. So let's call it d11, d22, and then d33. This equal to 3; this equal to 5; this equal to 7. And all the other elements, as you can see in here, they are zeros. So the concept of diagonal matrices is very simple; therefore, we will then go through the next example which is about orthogonal matrix. Now this is a concept that we haven't yet seen and we spoken about, so let's go read through this bit slowly. So an orthogonal matrix is a square matrix whose columns and rows are orthogonal unit vectors, so orthonormal vectors, and its transpose equals its inverse. So there are two parts of this elements. So firstly, it says that for the matrix to be orthogonal matrix it should be a square matrix, so square matrix, and then its columns and rows are orthogonal unit vectors, so columns and rows are orthogonal unit vectors, which means they need to be normalized, so normalized. So we have seen when forming this orthonormal basis that we had this process of this condition of orthogonality; the vectors had to be orthogonal and they had to have a length of one, which means that they they had to be normal. We can see exactly the same in here, so hence the name orthonormal vectors. So they are orthogonal and they are unit vectors, which means they are normalized. So then the final condition is added in here, which actually is not so much a condition but rather than a property, something that we can prove that once we have all this we can also say that if we are dealing with orthogonal matrix, then QT * Q, so the dot product of the transpose with that matrix Q is equal to the Q * QT is equal to I. Why? Because the QT is equal to the Q-1 because the transpose of that matrix Q is actually equal to its inverse, and given that we learned that the Q-1, so the inverse times Q is = to Q inverse * Q is equal to I, and given that here we are learning that QT is = to Q-1, we are then making use of this to claim this. So instead of minus ones that we are used to when we are dealing with inverses, here we have t, the transpose. So in this case, this orthogonal matrix that we have just learned about, this is the square matrix whose columns and rows are orthogonal and they are also normalized, meaning that we are dealing with QT; QT should be Q-1. It will look like this: so this Q1, you can see that here we got the first row; here we got the second row. And if we calculate the dot product between this row and this row, we can quickly see that we are getting a value of zero. So we can prove that they are actually orthogonal; those two rows. Let's go ahead and actually prove that. So let's call this R1; let's call this R2; this is row one and row two. And I will leave the column version, so Q1 * Q2, that's dot product, on you to prove that the columns are perpendicular. I will work with the rows, so R1 * R2 for me to prove that they are orthogonal; I need to prove that this equal to zero. Can we do that? Well, let's try. So 1/√2, 1/√2 multiplied by 1/√2 needs some bigger space in here; so 1/√2 and then -1/√2; that's how I can calculate the dot product between R1 and R2; R2 and R1. You can see that the elements in here are the same, and here the elements are also the same. So then this is equal to 1/√2 * 1/√2 - 1/√2 * 1/√2. I'm simply taking this minus and given that the dot product is basically plus and then this amount I'm just taking this and bringing up in here to avoid one more step given the space is quite limited. Now what do we see in here? This value is the same as this value, which means that this is equal to zero. And we know that the two vectors to be orthogonal, they need to have a dot product equal to zero. So here we have proven that dot product of R1 and R2 is equal to zero. So this proves that R1 and R2, so the two rows of this matrix, so R1 and R2 are orthogonal. We can also prove that the second criteria of orthonormal vectors is also satisfied in here. We can prove that when we look at the length of this vector and of this one, then they are of the unit one. So let's actually go ahead and do for one of them. So let's prove that for 1/√2 and then 1/√2; this is a vector that the length of it, this is let's say our first row, so this is R1, then the R1 length is equal to 1/√2^2 + 1/√2^2. This is equal to 1/2 + 1/2, and what is 1/2 + 1/2? It's equal to 1. So we have proven that the length of R1 is equal to 1. You can quickly and easily also compute that for the second row, and you will then also prove that the R2, the length of it is also one, which is then the second criteria which said that for the vectors to form this orthonormal bases, so to be orthonormal vectors, they also had to have a length of one, so they had to be unit vectors. In this case, then we can make use of the property that the Q2 transpose is equal to Q inverse, and this then results in Q2 transpose * Q2, Q2 which is equal to Q2 * Q2 transpose which is equal to the identity matrix, and specifically I2 because we are in the R2. So both this Q2 and the previous example, those are orthogonal matrices, and in here we have proven that the rows are indeed orthogonal, and we have also seen that the length of them are unit vectors, meaning that we have automatically got this. I will leave this one to you to do those proofs, so to ensure that the row one and row two are orthogonal, so they are perpendicular, which means that the product of their the dot product of these two vectors is equal to zero, and also that they are normalized, which means the length of them is equal to one, and this means that then this holds. You can actually even go ahead and practice the material that we are learning as part of the previous units by calculating the inverse of this matrix and checking that the inverse of this matrix is indeed equal to the transpose of the matrix, so that QT is equal to Q-1 because we learned how we can compute the inverse of a matrix because the inverse of a matrix was equal to 1/ the determinant of this matrix times and then the manipulated version of it which was in this case 0 0 and then we need to have here 1, so -1 * 1 and then 1 * -1, so we have to multiply this and this by minus, so one and then here minus one. So in this way you can also prove that this inverse is actually equal to the Q2 transpose because then you can prove that indeed and you can see for yourself that this formula is indeed true.

In this module, we are going to talk about matrix factorization. We are going to discuss the significance of matrix factorization. We are going to define matrix factorization. We are going also to discuss the common applications of matrix factorization across different fields, and then we are going to see detailed examples of matrix factorization. So let's talk about why matrix factorization matters. So matrix factorization techniques, they are essential for various reasons. They are used for simplifying matrix operations like solving linear systems or when we have these many matrices, but we want to to simplify these operations that we apply to these matrices and we want to solve the problem, then we can make this complex matrix operations more manageable and make these calculations more manageable by using matrix factorization techniques. We can also use matrix factorization directly to solve systems of linear equations efficiently. We can also use matrix factorization to perform eigenvalue decomposition, singular value decomposition or called SVD, and other operations which are crucial in machine learning and data analysis. So eigenvalues and eigenvectors, you might have heard already, they are part of also PCA, which is the principal component analysis, and this comes from fundamentals of statistics and the PCA is used as a dimensionality reduction technique, and in fact it's one of the most popular dimensionality reduction techniques that you will find in the industry, used in the data science, using data analytics, machine learning, even in the deep learning. So matrix factorization can also be used to reduce the computational complexity. By making use of this factorization, we can then simplify the process and also make it more efficient for computation, and it's especially important when we are dealing with this high-dimensional data. When we have many features or we have a very large model and complex model, then this matrix factorization technique can make a huge difference in our data processing process. So this techniques underpin many algorithms in numerical analysis, in optimizations, and beyond. So whenever it comes to machine learning or data science or many other fields, you will see this process and this term matrix factorization appearing a lot, even in the example of a streaming company, Netflix, which I'm sure that you are aware of. Netflix is using matrix factorization to build a recommender system, and matrix factorization usage in building recommender algorithm for personalized recommendations is actually one of the most popular applications of matrix factorization. Therefore, I wanted to specifically discuss this topic as part of our advanced linear algebra course. And some of the concepts might seem bit more complex than the ones that we have discussed as part of the previous units, but once we go through them step by step and I will give you all the details in all these examples, this entire process of this different matrix factorization techniques should become much more clear and straightforward. So we will be discussing not just one but multiple fundamental matrix factorization techniques, beside of talking the high level where they are used and how you can choose for what type of applications. So we are going to demystify this entire concept of matrix factorization, and we are going to start from high level then we are going to go into the deepest details. Let's now formally define the matrix factorization. So matrix factorization refers to decomposing a matrix into a product of two or more matrices, revealing its structure and simplifying further analysis. So what is this idea behind matrix factorization? The idea is that if we have a matrix A and we want to simplify our process of calculation or multiplication, anything that's related to this A, but this A in itself it contains these weird numbers or it is just too complex, you know it contains this ton of different numbers, you don't recognize whether the columns are linearly...

Independent; it's not very readable from the first view, and you just want to make your life easier when performing these calculations. Well, for that, you can make use of this matrix factorization to write this A in terms of some other matrices. Let's say U and I'm calling here randomly Q or T; so it's equal to, for instance, the dot product of these two matrices Q * T, where Q is much simpler and T is also much simpler. So those may contain vectors that are, um, for instance, this can be a diagonal matrix, or it can be a matrix, uh, with specific properties. When using those, you will feel much more comfortable, so it will be easier for you to use them in order to multiply, uh, with other matrices. It can be easier for you to solve this problem. But of course, if you are in the two-dimensional space, let's say you are in R² or in R³, then most likely it will be quite straightforward for you to use the A itself. But if you are in R¹⁰⁰ or R¹⁰⁰⁰, then of course this, uh, entire computation, they become super complex. It will be difficult to understand and compute these linear combinations, find out whether you are dealing with linearly independent columns, find out, um, the, um, null space, the column space, the basis of the null space and column space, and all these they might seem, uh, much more difficult when you are in high-dimensional space. For those cases, we can then make use of matrix factorization to make the entire process much more simplified and also more efficient; this entire calculation process.

Common types of matrix factorization include lower upper, uh, matrix factorization, or in short LU factorization; an infamous type of factorization which is called orthogonal triangular factorization; and then we have SVD, singular value decomposition; yet another infamous matrix factorization; and then finally, the eigen decomposition, also another super popular matrix factorization technique. So, uh, the QR, SVD, and eigen decomposition are, in fact, highly popular decomposition and matrix factorization techniques that you will see appearing in 90% of all the statistics-related and machine learning-related books. So this just comes to prove how important these concepts are when it comes to properly learning and mastering these more applied, uh, science-related fields like machine learning.

So if you want to go beyond the level of knowing algorithms, but rather than to also be able to edit the algorithms, tweak them, adjust them, be able to understand machine learning algorithms, deep learning algorithms, data science at its core, and in order to become a well-rounded professional, then this eigen decomposition, the singular value decomposition, and the QR matrix factorization techniques are techniques that you want to know and you want to understand at least at a higher level, such that you can easier grasp more advanced concepts that come from the applied sciences like machine learning and AI.

So let's first discuss at a high level what this QR decomposition is. So what the QR decomposition does is that it decomposes a matrix into an orthogonal matrix, which we are referring to by Q, and then an upper triangular matrix R. So in here you can see that we have these two different matrices. So we are basically saying A is equal to this product of these matrices Q and R, where the first one, this matrix Q, this one should be an orthogonal matrix. So this part is really important, and we have learned as part of the previous module the definition of an orthogonal matrix. We learned that the rows or columns they had to be orthogonal to each other, and we also learned that they need to have a length of one; they need to, um, be normalized. And we learned that this means that the transpose of those matrices is equal to the inverse of the matrices. So this was just the last part of the previous module, and this is exactly what this matrix Q is about. So we are saying that we will decompose A into these two matrices as a product of these two matrices Q and R, one of which, this matrix Q, should, uh, be, uh, an orthogonal matrix, which means the rows and the columns they should be orthogonal to each other. So their, uh, dot product, each of them should be equal to zero, and they need to be normalized, so the length of them should be one for each of those rows and vectors. And then the second part of this QR decomposition is this matrix R, which says that the matrix R should be an upper triangular matrix. And what is the definition of upper triangular? Well, in this case, you can think of this R as this matrix where we have here all zeros, and then here on the diagonal you have numbers, nonzero numbers, and then here, let's say 1, 2, 3, 4, 5, and then here in the upper part you will also have numbers. So unlike in the lower part of this matrix R where you will have zeros in here, you will have also nonzero numbers, numbers, let's say seven, uh, ten, uh, eight, and I'm just writing these numbers randomly. So of course, in a real case when we have this matrix A and we go through this process of QR decomposition, of course, we will have an appropriate Q and appropriate R where these numbers will be different and they will be specific numbers that we will be calculating. But the idea is that we need to get this upper triangular matrix R for this calculation to make sense.

So we will then be using this QR decomposition for solving linear, um, linear least squares problems, for instance, which is part of the linear regression too, because linear regression from machine learning and from statistics, uh, it is based on the least squares technique; the estimation technique that we are using for linear regression in machine learning, um, to solve this linear regression problem is called Ordinary Least Squares. So the algorithm is based on this idea of least squares, which is trying to minimize a squared, uh, residuals of the model, and that can be done by using this idea of QR decomposition. So it helps us to provide numerically stable solutions for this type of problem too, and QR decomposition is used extensively in signal processing and statistical analysis.

Let's now briefly talk about the LU decomposition. LU decomposition decomposes a matrix into a lower triangular matrix. So this is the opposite of what we had before; we can have an upper triangular matrix like we had in the QR decomposition; we can also have a lower triangular matrix. So you might have already guessed how it will look like. I won't go into that very soon. In the QR decomposition example, you will see the idea of the upper triangular; I will also show the idea of a lower triangular. So the decomposition, in case of LU, is done by decomposing a matrix into a lower triangular matrix L and an upper triangular matrix U. So basically, the difference between the QR decomposition and LU decomposition is that in the QR decomposition we are decomposing a matrix into an orthogonal matrix and an upper triangular matrix, while in case of LU decomposition we are decomposing a matrix into a lower triangular matrix and an upper triangular matrix. So here you can see that we no longer have this idea of an orthogonal matrix, but instead of that we are talking about a lower triangular matrix. So in that aspect, uh, LU decomposition is different from QR decomposition. So what the LU decomposition does is that it facilitates the solving of linear equations and matrix inversions. It is common in engineering and in physical sciences for systems of these linear equations to be solved by using LU decomposition, and in fact, if you are learning quantum, uh, mechanics, that this LU decomposition can definitely help you to better understand many concepts. But if your target fields are machine learning, deep learning, or artificial intelligence, then for those using QR decomposition will be, uh, much more often a case than using this LU decomposition.

Let's now talk about the singular value decomposition. So what the SVD does is that it decomposes a matrix into three matrices. So the first one is the orthogonal matrix U, the second one is a diagonal matrix Σ, and then the third one is this V, which is the conjugate transpose of an orthogonal matrix. So for now, this might seem a bit complex, and you can see that unlike the QR or LU decomposition where we got just, uh, two, uh, matrices as a result of our decomposition, in case of SVD we got three matrices, like the name suggests, two, by the way, so three parts, and this might seem a bit complex, but we are going to go through this process step by step, and I'm going to provide you a detailed example such that this will all make sense. But for now, let's focus at the high-level usage of SVD. So singular value decomposition is one of the most popular decomposition techniques, and it is also directly used as part of machine learning algorithms to form, uh, those machine learning algorithms. It is also used in data compression, in noise reduction. So when we are trying to clean our data and remove the noise by using SVD, because SVD can help us to identify those outliers and then remove them from the data by performing noise reduction, and it is also used in principal component analysis, the PCA, the, uh, same dimensionality reduction technique that I just, uh, mentioned related to the eigen decomposition, because SVD and the eigen decomposition are highly related to each other. So this SVD is used as part of this PCA algorithm, and PCA is the most popular, the infamous dimensionality reduction technique that is used both in advanced statistical studies in statistics in general, also in finance, and is also used as part of many machine learning and deep learning applications. So knowing PCA is a must if you want to get into, uh, data analytics or data science, machine learning, and AI, but also, uh, it helped you, it will help you also to, uh, understand, uh, many other concepts when it comes to these fields. So the SVD provides insight into the structure, but also the rank of the matrix. So we are going to see this as part of our example too.

Let's now also briefly talk about the eigen decomposition. Eigen decomposition, which is highly related to this concept of eigenvalues and eigenvectors, it decomposes a matrix into these eigenvalues and eigenvectors, which then shows the matrix's fundamental properties which are related to this idea of correlation; what kind of information does this, uh, matrix contain; what is the variation; in what direction is the variation the largest; and this eigen decomposition, which is then related to also this idea of SVD, and in general this dimensions and correlations is critical for understanding linear transformations, the stability analysis, and systems of differential equations. But beside this, uh, mathematical side of, uh, concepts and understanding these mathematical topics, the eigen decomposition is also the basis for many algorithms in numerical linear algebra, uh, but also many applied linear algebra topics like in data science, in machine learning, and it's used heavily in artificial intelligence for feature extraction, for dimensionality reduction, related again to the concept of PCA, because PCA is based entirely on this concept of eigen decomposition. PCA is the direct result of computing the eigenvalues and eigenvectors. So without knowing what are eigenvalues and eigenvectors, you cannot perform PCA, because the first step of the PCA is the computation of the eigenvalues and eigenvectors, and then using different rules, which we are referring to, uh, as the elbow rule or Kaiser rule, we can then use these eigenvalues and eigenvectors to understand what are the features in our data that contain the most variation, so the most information, and then we can use that in order to understand what are the most important features in our data and reduce the dimension of our model by selecting these most important features. Because what PCA basically does is that it uses these eigenvalues and eigenvectors to understand how we can, uh, create a linear combination out of our features and understand the amount of those linear combinations that contain the most variation, and then select those and, uh, select the largest amount of information in the data, and then skip and drop those uninformative linear combinations, while still keeping the most information. And this definition of the most will then be decided by these differentials. This is just higher-level insight, background information on what you can expect when you are talking about applying this highly technical linear algebra concept of eigen decomposition into an applied science field like data science or machine learning or AI. But we will see this later, and I'll also make comments regarding this. And though PCA won't be discussed as part of this course, because here we are talking about linear algebra, but PCA is part of the fundamentals statistics course, and in there we are no longer providing all these different details on how you can, uh, perform this eigen decomposition. Therefore, knowing how to perform eigen decomposition will then set you for success to actually understand the mathematics behind the statistical concepts like PCA and also later on understand how you can use that PCA in machine learning concepts and in AI concepts like autoencoders, and how you can relate your, um, autoencoders to this concept of PCA, how they are related, what are their commonalities and what are their differences.

So everything is about the choice and choosing the right tool for your problem. When it comes to the decomposition tools, matrix factorization tools, we have seen that there are many options, and the question is which one should we pick, in what cases? So choosing the right tool is really important when it comes to these different matrix factorization techniques, because there are many choices, and each of them they can be used for different sorts of problems. So therefore, in order to understand which one you need to pick in what kind of cases, what kind of requirements you have, and what kind of problem you are trying to solve, that in those cases you will need to have this knowledge that you will learn as part of this course in order to make that right choice of the tool. So the choice among QR decomposition, the LU decomposition, the SVD, and eigen decomposition, it really depends on your specific problem's requirements and the data characteristics. So are you dealing with complex data? Are you dealing with simple data with low dimensions? What is the goal that you, uh, want to, uh, achieve? What is the problem that you are trying to solve? Is it to, uh, reduce the dimension of your feature space? Is it to solve a problem with linear equations? Is it to solve a quantum mechanics problem, or is it to, um, incorporate this as part of your machine learning algorithm? So QR and LU decompositions are usually preferred for solving linear systems, while SVD and eigen decompositions they help us for deeper insights when it comes to the data and what kind of information it contains, how we can reduce the dimension of the data, or how we can use it as part of a machine learning algorithm for, uh, noise reduction, identifying outliers, etc. So, um, this type of algorithm, like SVD and eigen decomposition, it helps us to also, uh, intuitively using geometry and our knowledge of geometry to, um, visualize the data. For instance, the PCA helps us to visualize this high-dimensional data using just a couple of principal components. Let's say we have ten features in our model, so we have a dimension of ten; we are in R¹⁰, but we want to visualize our data by using PCA. We can then reduce the dimension and come up with, uh, three principal components, which are a linear combination of our original ten vectors, and then we can use the three principal components to visualize our data in 3D, and this basically helps us to geometrically visualize our data and then make presentations, make much more sense of our story, so to do, uh, storytelling for our data and, uh, much more. And these two, uh, models and tools they are invaluable when it comes to, uh, applications in machine learning, in deep learning, in data science, and artificial intelligence.

Matrix factorization techniques they are super important when it comes to computational mathematics; they are also directly affecting data science, AI, and many other algorithms. So they are not only important in terms of the problem that they are trying to solve, but also in order to make the computation process, so when coding in Python or in other programming languages, to make that process much more efficient. They also help us to, uh, make these computations efficient and provide insights into different properties that we have in our data. As part of this course, we are not only going to discuss one, but actually three of these four decomposition techniques and these matrix factorization techniques in detail. We are going to talk about the QR decomposition; we are going to not just discuss it, but, uh, also to learn it step by step, and we are going to do a detailed example with all the steps involved such that you will feel confident doing a QR decomposition all by yourself. Then we are also going to do an SVD decomposition as well as eigen decomposition, and then we are again going to discuss them in terms of their mathematical formulation, the definition, but also the application, step-by-step process, and a detailed example such that you can conduct each of those matrix factorization techniques and these decomposition techniques by yourself, manually doing all these calculations. This understanding and these examples and these concepts will help you to not just be able to formulate what these techniques are about, but really and truly understand and then use them later on, whether when doing your own research, writing scientific papers, or tweaking the algorithm all by yourself when inventing new algorithms. I won't be discussing this LU decomposition technique because we already know, uh, that the QR and LU they are both used for similar types of problems; therefore, to save us time, I have selected carefully the, uh, most important decomposition techniques and matrix factorization techniques that you will most likely be dealing with, will be dealing with in your future career in applied sciences.