📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Linear Algebra Course – Mathematics for Machine Learning and Generative AI

freeCodeCamp.org6:05:12

Transcription

Ready to learn the mathematics behind machine learning, generative AI, or cutting-edge tech? Then you are in the right place. Today, we are going to learn linear algebra like never before.

In this crash course (in 6 hours), we are going to learn the basics and fundamental concepts in linear algebra—a fundamental part of mathematics when it comes to many applied sciences. Our crash course in fundamentals to linear algebra will be a great starting point for anyone who wants to refresh their algebra and knowledge, or who wants to start with the fundamentals of mathematics for AI, data science, machine learning, or generative AI.

So, if you are looking for that one course to refresh your memory on those fundamental concepts and be able to create your own algorithms and create your own AI dream products, then you are in the right place. And if you're really serious about your education and career and backing up your mathematical foundations, then you can also check our 20+ hour comprehensive fundamentals to linear algebra course, which is also part of our mathematics boot camp. So check it out! And you can get this comprehensive 25+ hour course along with the certification in linear algebra.

Here's what we are going to cover as part of our fundamentals to linear algebra crash course. We are going to start with the introduction to linear algebra. We're going to talk about the prerequisites—what exactly you need in order to get into linear algebra, whether as a student (as part of your undergrad or graduate studies) or as a working professional who wants to learn the mathematics in a good way in order to become a well-rounded professional (whether it is for your future in quantum physics, data science, machine learning, or artificial intelligence).

So we are going to learn the basics of vectors. This is the first section that we are going to cover here. We are going to learn the vectors, their basic definitions, and the theory, but also the examples. In this course, we are not just going to scratch the surface; we are going to dive deep. The idea is that, even for those basic concepts, we are going to learn not just the theory and the basic definitions, but we are also going to implement them in different examples. We are going to manually solve different examples by hand in order to truly master and practice this topic. Beside this, we're also going to talk about the applications of them across different domains, including in data science and AI. The idea is that those basic concepts are your starting point in the future of linear algebra. Because to truly understand this topic and to truly understand this entire linear algebra field (at least the course, the academic level course that you are used to seeing as part of the different advanced mathematical studies), you will need to dedicate more than just 6 hours.

So if you are serious about your education and career and you want to learn mathematics from the ground up, then check also our 25+ hour complete fundamentals to linear algebra course to truly master the entire linear algebra in a very economic and efficient way, which contains an entire semester-equivalent material (theory, practice, and applications).

Here's what we are going to cover as part of our crash course. So we will start with the introduction to linear algebra. We will talk about this concept of linear algebra. We will see what are the prerequisites in order to even understand and start with this course. We're going to refresh our memory with the different fundamental concepts here in this introduction to linear algebra section. We are going to quickly cover the prerequisites that you must know in order to get started with linear algebra. Here, I'm also going to refresh our memory on the fundamental concepts like real numbers and vector spaces, the norm of a vector, the Cartesian coordinate system, and also the angles (trigonometry), the norm versus Euclidean distance, Pythagorean theorem, etc.

So, after this, we are going to finally kickstart dealing with linear algebra concepts. First up, we are going to talk about the basic vector spaces. Here we are going to cover the definition of the vectors—what are vectors, the basic representation of it, but also the indexing of it and how we can apply different simple operations like vector addition, vector subtraction, and also scalar multiplication. After this, in the next section, we are going to talk about the matrices and the basic operations. We are going to truly understand the matrices, the basic definition of them. We are going to see many examples of matrices and also, in practice, what are the basic operations that we can apply to matrices, including the additions, the scalar multiplication, and subtractions. After this, we are going to look into the dot product and the vector length. In this section, we are going to talk about the idea of dot product; what is the difference between dot product and inner product; understand the fundamentals of that product and how they are calculated. We're going to complete many examples, as this is a very important topic, especially for data science, machine learning, and AI. And once we are done with this, we will also look into this concept of the length of a vector and how this is related to, for instance, cosine rule and these different inequalities when it comes to the vector spaces. After this, we are going to look into the basic linear systems; how we can solve a linear system by using the infamous Gaussian elimination. Here we are going to talk about this basic concept of solving linear equations with many unknowns using the concept of matrices. We are going to see detailed examples where, manually, step by step, by hand, we are going to derive the entire process using Gaussian elimination; how we can solve the system of linear equations, first getting the augmented matrix, and then this concept of reduced row echelon form (RREF), so row echelon form and reduced row echelon form; and we are going to look into two different scenarios to ensure that we understand this intuition behind solving these linear systems in various types of cases.

Now we are going to finish off the crash course with a preview of the advanced concepts. We are going to see the special matrices; we are going to see symmetric matrices, diagonal matrices, and then the idea of n matrices, W matrices, how they are used and why they are used commonly in programming; and then we are going to see a glimpse of how linear algebra can be applied across real-world scenarios.

In this crash course, we are not just going to cover the quick definitions, but we are going to dive deep; we are going to truly try to understand those concepts, those geometric intuition and interpretation behind those concepts, and how they are related towards many applied sciences. To truly understand how each of those concepts are actually implemented in the real world (like in the fields of data science, statistics, and AI), this course is perfectly tailored for working professionals who need to sharpen their math skills, who want to quickly and efficiently refresh their memory (whether it is before their interviews) or who want to learn machine learning and statistics and AI, but they don't have the mathematical background. Instead of just going and studying at university like I did and spending many years to do an undergraduate studies in mathematics or statistics, instead you can just follow this course to quickly refresh your memory. But if you truly want to understand all the concepts (which is equivalent to one semester of linear algebra course), then definitely check out our 25+ hour fundamentals to linear algebra course as part of our mathematics boot camp, because you can't master linear algebra in just 6 hours.

If you are a student, this course will also be a good way for you to refresh your memory if you are at the end of your deadline and you are about to complete your linear algebra exam. And if you want to learn everything from scratch and you have the time, then you can also check our 25+ hour course to learn the linear algebra fully and to ace your exams. This crash course is just the beginning for you; in this way, you are laying a solid foundation to become a well-rounded professional and to truly understand this topic of linear algebra.

So if you're ready, I'm really excited. So let's get started! Welcome to the mathematics for data science and AI roadmap in 2024. Today we are going to discuss the roadmap behind a topic that powers many algorithms: linear algebra. We are going to talk about all the concepts and the topics that you must know in order to master linear algebra, which is the backbone of data science, machine learning, generative AI, and many other topics and cutting-edge tech. Whether you are aspiring to become a data scientist or AI professional, this is your starting point to learn the mathematics of data—the fundamentals of linear algebra.

If you want to become a good professional, a well-rounded data scientist, and master the art of machine learning, generative AI, and many other cutting-edge tech domains, then your starting point should be to learn the mathematics—to truly understand these different mathematical concepts that power the algorithms behind the scenes, including many infamous algorithms like the GPT series, or the BERT, or Transformers. Linear algebra is the mathematics of data; it's about the vectors, matrices, projections, and how to solve a system of linear equations and how to decompose your matrix into many parts to simplify the equations and solving of the linear systems. Think about matrix factorization that is used as part of the recommended systems, including the ones in Netflix (that streaming company) to recommend you movies when you go to the net. So linear algebra is one of the most fundamental and must-to-know topics when it comes to mathematics for your data science AI journey.

So let's now dive into the roadmap behind linear algebra and how you can learn linear algebra in 2024. I'm going to make use of the curriculum and the roadmap that we have carefully drafted as part of our 25+ hour fundamental linear algebra course, which is also part of our mathematics boot camp. So I'm going to make use of that roadmap, and I'm going to also make use of the applications of linear algebra in order to showcase you all these different topics as part of the linear algebra roadmap of 2024. So let's dive in!

First, we need to start with the linear algebra introduction. You will need to understand the concepts behind linear algebra, where it is positioned in this entire field of mathematics, and why you exactly need linear algebra to study. So, first, you need to learn the prerequisites and all these must-know topics that you need in order to be able to understand linear algebra. Think about the real numbers, the vector space, this concept of the norm of a vector, the length of a vector, the Cartesian coordinate system (which is usually covered as part of the pre-algebra or algebra courses), and also the angles and trigonometry. Think about the cosine, the sine, cosine rule, the sine rule, the tangent function, how it looks like when you visualize it, the geometric interpretation behind it, and also this concept of Euclidean distance and how Euclidean distance is different from the norm. Finally, you also need to understand this concept of the Pythagorean theorem and the orthogonality.

So, when it comes to linear algebra, one thing to keep in mind is that linear algebra is usually a course that is part of the undergrad studies (like computer science, econometrics, advanced mathematics, advanced statistics), and usually the students of these studies, in the second or third year of their bachelor's, they learn linear algebra. And this roadmap that I'm providing to you is based on academic-level plus the industry insight curriculum, which means that the concept is quite technical. So if you want to truly understand linear algebra and you are serious about your career and your education, then this roadmap, on one hand, might seem a bit technical and intimidating, but on the other hand, you need to understand that students spend about a semester on this entire roadmap. But if you understand these concepts, this will open many doors for you. And instead of just using the libraries of people who have created, for example, large language models like the GPTs, or the BERT, or the mixture AI models, instead of just using those libraries and pre-trained models, you can create your own models. And I think that's a good enough motivation for us to truly understand and invest the time and motivation to master the art of linear algebra.

So once you have these prerequisites (that usually come from pre-algebra or courses like Calculus 1 and Calculus 2) to learn these different concepts, you are ready to go onto the first journey and first section behind linear algebra, which is to learn the vectors and operations. Be prepared to learn the foundations of vectors, the special vectors and operations. And let me actually go ahead and get more detailed curriculum for you.

So, start to understand vectors and operations—what are the foundations of vectors, the fundamentals of linear algebra, the scalars and the vectors, this representation of the vectors in terms of the magnitude and the direction, and also understand this vector notation, the indexing of the vectors. First up, you can start with learning vectors and operations. So here you need to understand these scalars, the vectors, the representation of the vectors like the magnitude and direction, what is a common notation of the vectors in the industry, as well as this indexing in vectors. So understand what the indexing behind vectors are, because then you need to implement this in programming languages like Python. Then, once you are clear on those topics, then you can get into the special vectors and operations. So understand what are the zero vectors and the unit vectors, the sparsity vectors, the vectors in higher dimensions, the vector addition and subtraction, as well as the vector multiplication—scalar multiplication; how you can calculate the multiplication behind the scalar and a vector; also learn the different properties of the vectors and the operations of these vectors. So once you are done with that, get into the advanced vector concepts. So learn the linear combination concept and try to understand this concept of unit vectors, the span of a vector (what is this definition and the application of it), the concept of linear independence (is really important; understand this concept and also be able to analyze and see from multiple vectors whether they are linearly independent or not), look into this concept of scalar multiplication, as well as the application of scalar-vector multiplication, and look into this concept of a length of a vector and a dot product, as well as what is the difference between the dot product and the inner product. After this, get into the topic of dot product, the Cauchy-Schwarz inequality, and its application. So understand the concept of dot product, the inner product, what are the differences between the two, what is the properties of dot product, calculate dot product by hand on many examples, and also learn this concept of Cauchy-Schwarz. Cauchy-Schwarz is an infamous inequality theorem that comes from linear algebra, and understanding its theory but also intuition—maybe you can even derive the proof of it using the cosine law and how that relates to this concept of the norm of a vector and when it is that the norm of a vector is equal to zero. So this completes then the vector operation section. In here, this is the high-level overview of the topics that you must know in order to learn the vectors and the operations. And once you are done with this, then you need to get into the next section and the domain in linear algebra, which is the matrices and solving linear systems. Here I'm talking about foundations of linear systems and matrices, the introduction to matrices, the theory and the intuition behind it, core matrix operations, Gaussian reduction, null space, column space, rank, and the full rank.

So here think about the general linear systems, like the coefficient labeling, homogeneous versus non-homogeneous systems, and what is this definition of matrices, the notation of it, as well as this concept of the rows, the columns, the dimensions of matrices, the identity matrix, diagonal matrix, other special types of matrices like ones matrix, the zero matrix, but also learn the core matrix operations and practice with them, like matrix operations of addition, subtraction, scalar multiplication, and also the multiplication behind multiple matrices. Once you are done with this, in terms of the solving of linear systems, think about matrix representation of linear systems, the solving of linear systems using the infamous Gaussian elimination, like understanding this augmented coefficient matrix, the row echelon form, the reduced row echelon form, the identity matrix and the reduced row echelon form. Look into the detailed examples; implement this in practice with multiple examples, step by step, by manually deriving the solution to the system, as well as how you can calculate a null space, a column space, and the basis of these different spaces on an actual example. Once you are done with this, learn this concept of the rank and in what case you can say that the matrix has a full rank.

Next up is the section of linear transformation and matrices. So learn the algebraic laws of matrices, their proofs, the determinant, the concept of the transpose, and the inverses of matrices. And here this is exactly what I mean in more detail. So when it comes to the algebraic laws for matrices, learn about the commutative law for matrices, associative law, distributive law, the scalar multiplication law for matrices, and look into the proof of these laws as well as the examples. After this, get into the determinants and their properties; look into the theory behind determinants, why we use determinant, how we can calculate the determinant on an example, as well as what are the properties of determinants and the geometric intuition behind it. So when it comes to the matrix inverses and identity matrix, look into the definition, the theory behind it, also this non-singularity of matrices and what is this definition of, and the formula when it comes to the inverse of a 2x2 matrix and how you can go from there to larger matrices, let's say 3x3 matrices. Looking to the detailed examples, calculate them inverses of the 2x2 matrices in detailed examples, and also do the same for the 3x3 matrices; so how you can use this idea of conjugates and by changing these different signs to calculate the inverse of a 3x3 matrix. After this, get into the applications of inverses of matrices—why we need them, what is the intuition behind it, and where do we use it, even when we are dealing with the very basic machine learning algorithm like linear regression. We use actually the transpose of matrices, the inverse of matrices, the multiplication behind matrices to find the solution to the linear system behind the linear regression. So this is a very important concept that I recommend you to learn before getting into machine learning. Next up is the transpose of matrices; learn the properties of the transpose of matrices, the theory behind it, but also the implementation and the relationship behind the dot product.

Once you are done with this, the next step is to get into more advanced section, which is the final section when it comes to mastering fundamentals of linear algebra, which is these different advanced linear algebra topics. Think about vector spaces, the projections of vectors, the Gram-Schmidt process, the matrix factorization, and look into a couple of examples of matrix factorization techniques that are widely used across the industry, including the QR decomposition, the eigen decomposition (which includes the calculation of eigenvalues and eigenvectors—very popular topic as part of data science and statistical modeling, also AI), and get into the singular value decomposition (which is the SVD).

Let's dive into the detailed details of this section. When it comes to the vector spaces, projections, and the Gram-Schmidt process, look into the definition of it, the step-by-step process behind it, the projection formula, and the step-by-step example behind the projection in order to truly understand how you can calculate the projection of a vector, and also look into the geometric interpretation and the intuition behind vector projections. Learn the theory but also the implementation of the dot product of vectors and the dot product of a vector with itself. Look into the relationship between the dot product, the vector projection of vectors, and then look into the concept of orthonormal basis; understand the orthogonality, normalization, and look into the infamous Gram-Schmidt process; understand how you can use the idea of dot products, but also the projections in order to conduct this Gram-Schmidt process step by step. Once you are done with this, then look into the application of Gram-Schmidt process and the orthonormal bases. As part of these special matrices and their properties, I would recommend you to learn the special matrices like the diagonal matrix, the symmetric matrix, the orthogonal matrix, and the different applications and examples of them step by step.

You are done with this. Learn this concept of Matrix factorization. Matrix factorization is a very important topic when it comes to data science, machine learning, but also AI. Uh, the matrix factorization: different examples and techniques that um are falling under this umbrella of matrix factorization techniques that are used in both in machine learning, but also in deep learning and in generative AI. So if you want to become a true uh well-rounded professional in this field, then definitely be prepared to learn matrix authorization techniques, dive into its applications, understand the different examples of what kind of matrix alization techniques are used in what uh algorithms. This will help you to get more excited and warm up for the next sections, which are detailed examples, and you truly understand at least a couple of matrix factorization techniques.

So first up, uh learn, for example, the QR decomposition, which is a matrix factorization technique. Learn the theory behind it, the step-by-step process behind QR decomposition, how you can calculate the um Q's and then the vectors of R. So construct the matrix of Q and then matrix R, and then uh decompose your matrix A into uh Q and then R matrices in order to decompose your entire matrix. Look into a detailed example, the calculation of it uh with a step-by-step manual process, and look into its practical consideration and application in solving linear systems.

After you are done with this, look into the concept of eigenvalue decomposition. How you can calculate the eigenvalues and eigenvectors. What is the theory, but also the practical implications of them? Where are eigenvalues and eigenvectors used? For example, in the infamous principal component analysis or PCA. Then look into this concept of eigen decomposition. Look into detailed examples and manually calculate the eigenvalues and eigenvectors, as well as the eigenvalue decomposition.

So once you are done with this and you have looked into the applications of eigen bi and eigenvectors, the calculation of it as well as the theory, getting into the final topic which is another matrix factorization technique, an infamous one which is the singular value decomposition, the SVD. Look into the theory behind the SVD, the left singular matrices, the right singular matrices, this uh theory behind them, but also the practical example behind them. Uh, what is this idea behind um singular value decomposition? How you can construct the matrix of S, the matrix of V, and the matrix of D, and how you can decompose a matrix into the three different matrices by constructing them separately.

Once you are done with this, look into a detailed example. This might take a bit, let's say an half hour or 1 hour in the first go to calculate this different matrices by hand, but trust me it is worth it because you will then truly understand those concepts and you will understand how and why those uh matrix factorization techniques are used and applied in data science and artificial intelligence.

Welcome to the course on the fundamentals of linear algebra. My name is D Vasan, and today we are going to start with some basic concepts that are important for understanding linear algebra. Linear algebra is one of the most applicable areas of mathematics. It is used by pure mathematicians that you will see in universities doing research, publishing research papers, but also by the mathematically trained scientists of all disciplines. This is really one of those areas in mathematics that you will see time and time again appearing in your professional life.

If you want to become a job-ready uh data scientist or you want to do some hands-on machine learning, deep learning and AI stuff, but also linear algebra is used in cryptology, it is used in cyber security and in many other areas of computer science and artificial intelligence. So if you want to become this well-rounded professional, you want to go beyond using libraries and you want to truly understand the uh mathematics and the technical side of these different machine learning algorithms from very basic ones like linear regression to most complex ones coming from deep learning like architectures in neural networks, how the optimization algorithms work, how the gradient descent works and all these other uh different methods and models, then you are in the right place because you must know linear algebra such that you will understand these different concepts from very basic ones to most advanced ones in data science, machine learning, deep learning, artificial intelligence, data analytics, but also in many other applied science disciplines.

So before starting this comprehensive course that will give you everything that you need to know about linear algebra, first I'm going to tell you what we assume that you already know because linear algebra, it comes from about third uh year of Bachelor's um of different uh highly technical studies and um here um we are assuming that you already know certain concepts. So uh to ensure that this course stays really on the topic of linear algebra and that you uh understand all these concepts really well, for that we need to uh be able to know different topics. So before we dive into these concepts, uh let's familiarize ourselves with the basic prerequisites and notations used throughout this course, and you will really need to know this in order to understand this concepts really well such that instead of memorizing you will actually just hear me once or maybe twice and then every time you hear later on or you see it in the papers or in some algorithms you will recognize um this is something that we already learned.

So uh some key prerequisites overview is here. Um first of all, to fully grasp the upcoming material you should be familiar with some basic concepts like real numbers, vector spaces. So you don't need to know this idea of vectors, though you uh already most likely are familiar with this given that you know how to plot different uh lines, you know the idea of x's and y's and how to plot these different graphs. But um here we are going to touch base on this every time when we come close to this concepts. I will refresh you uh your memory and we will go through this numbers, the idea of norms and distance measures because when it comes to the vectors, when it comes to the magnitude and all these different uh topics that we are going to discuss as part of linear algebra, knowing the what norm is and um what is the definition of distance, what is the length between uh two points when we plot it in the two-dimensional space or three-dimensional space, those are all very basic concepts that usually you see as part of a basic pre-algebra or just a common algebra courses and um lessons.

In order to truly understand what linear algebra is about, to understand this direction of vectors, the angle and then um the uh dimensionality reduction, how linear algebra is applied for instance in different algorithms in machine learning, deep learning, data science, statistics, you really need to understand this Cartesian coordinate system. So uh this is not only important for linear algebra, but I assume you already know it given that you have passed those um uh other courses like uh calculus or usually they are covered as part of pre-algebra or algebra. So the Cartesian coordinate system, I mean here understanding uh what is for instance the the common um a description of them, for instance when you when we write like X and then Y on the vertical axis and then we can uh we have here zero and um then uh we can always plot this different plots. You know, we we have a clear understanding what is this um Y is equal to X line. We understand how by knowing certain points we can plot different plots, for instance that this is the Y is equal to X line that here it means that if we have here one then this is just one, two, this is two. So we understand when we have the function of the line and we have a certain value that is our y coordinate or x coordinate, then the corresponding uh coordinates can be found. Then um you also need to know um some basic things that I just didn't mention uh right now. So for instance that the numbers here can be like 1, 2, 3 up to infinity. So you understand this concepts of infinity and then here the same uh story. Then here we have minus one, you know, minus two uh and then this is then used later on and we will be uh touch basing this one, we will be describing our vectors and how uh we can visualize our vectors in a two-dimensional space like we have here because this is two-dimensional, so we have X and Y, but we can also of course visualize it in three-dimensional etc. So this idea of basic coordinate system is really important, um usually covered as part of algebra, if not pre-algebra.

Then we have basic trigonometry, which means that you need to have a clear understanding what sine is, what cosine is, what tangent is and their reciprocals. And here I mean uh that you know for instance um what is cosine function, what is sine function, um you know that you have an understanding for instance that um uh what is this line, you know um whether it's a sine line or cosine line. You have also an understanding what this Pi is. Um one thing that I didn't mention but it it just goes um around all these topics, some basic things that you understand what is X, what is Y, why we uh use them and this idea of uh variables uh and also uh you need to understand this idea of square uh or you know a 90° uh angle and then uh Pythagoras theorem. Here we have the same, so what is this relationship between different sides of a triangle that is a very unique triangle and that has one of the uh angles as 90°. Um and uh this idea of um you know the sides, how this relates to the sine, cosine, tangent, cotangent um and also um how the Pythagorean theorem applies when we have uh uh triangle but it is no longer with an angle that is 90°, what is the sum of all the angles of a triangle? So those are basic stuff that are commonly covered as part of uh trigonometric uh lessons or part of general geometry.

When it comes to this R, so as part of real numbers and vector spaces, R represents the set of all real numbers. So you can be dealing with for instance integers like 1, 2, 3. This can also this will also cover all the negative numbers like minus one, minus two, minus three, but also the floating numbers like 1.223 and all the other numbers that you can think of. Those are the set of all real numbers. So this is in one-dimensional space, right? So you can see that I'm writing just one number, you know, two, three and other numeric numbers. Then we have the idea of R2, R3 up to RN, where all these numbers they represent in this case the N, it represents the N-dimensional Euclidean space. So when it comes to this idea of N-dimensional numbers, so for instance R2 here we just mean 2D plane. So I'm pretty sure you are familiar with this idea of for instance x-axis and y-axis. Here we are dealing with two-dimensional plane. So for every point that we can find here we can describe them by assigning them a value X, so coordinate X and a coordinate Y. That's exactly what we mean by saying that the number can be represented in a 2D plane. So here we are dealing with this two-dimensional space, this is our two-dimensional Euclidean space and every number in here that is part of this R2 can can be pictured here, can be represented in this visualization. So for instance if I have this number and let's assume that the value on the x-axis is two and we can see here that the corresponding Y is zero, I can describe this number which I will call a, I can describe this by writing down first the x coordinate which is two and then the y-coordinate which is zero. So I'm then saying that a, which is a point with x coordinate 2 and y coordinate 0, it is part of my R2 and it's part of my two-dimensional Euclidean space.

When it comes to R3, a similar thing we can do with that, only in that case we need not just x-axis and y-axis, but we need to add our third dimension. So here for instance when it comes to the R3, then we need to do y-axis, we need to have x-axis, but also we need to have some z-axis, so such that every time every point in the space we can then describe by x, y and z coordinates. So if we write it in terms of the vector, something that we will see very soon as part of our first unit of this course, we will then need to represent every number in this three-dimensional Euclidean space by writing down first the x coordinate, let's say one, and then y coordinate, let's say another one, and then z coordinate which is one or even better, even easier, let's use 0, 0, 0, which means that we are dealing with this initial number which is the center of this three-dimensional Euclidean space. When it comes to the N-dimensional or higher-dimensional spaces, it's much harder to visualize, therefore usually when it comes to visualizations we do usually we usually only visualize the one-dimensional, two-dimensional and three-dimensional spaces. Above then it just no longer does make sense to visualize it, but we definitely deal with them and they are part of applied linear algebra. So understanding the spaces is very important for analyzing vectors for their interactions and this holds not just for this two-dimensional and three-dimensional but really for multi-dimensional space.

Let's now quickly define this idea of norm. So the norm of a vector, denoted by this uh uh V, which you can see kind of like similar to the absolute value from pre-algebra, you can see here that we have this double straight lines like from absolute value, then we have the name of the vector or the variable name that we are assigning to our vector and then you might notice is here on the top of this this arrow. This basically says that we are dealing not with just a variable but really we are dealing with a vector. This is really important because you can see that there makes a huge difference if we have for instance just V or V1 I have to say or just V. Those are really important and things that you need to keep in mind when it comes to linear algebra and trying to differentiate vectors from a point. You will notice that when it comes to norm we can uh represent it either by this notation or this. Usually it's a common um notation uh in machine learning or in data science um with this uh two bars and um when we do this we automatically also know L2 norm and this is something very common and uh usually used as part of ridge regression which is an application of um linear algebra uh and it's used in uh regularization. So we are regularizing our machine learning algorithms. So when you get into machine learning you will see time and time again this um notation. So uh next time when you see this then you know automatically that you are dealing with L2 norm and L2 norm which is also used a lot in machine learning, it is referring to the usage of L2 norm to uh in the uh ridge regression and regression or L2 regularization is a very popular regularization techniques as part of machine learning. So right now even you can see this uh intersection of linear algebra or um this uh idea of norms in machine learning. The norm of this vector v is equal to square root and then V1 squared plus V2 squared plus and all this in between numbers plus Vn squared. So here basically it means take square root of V1 squared then V2 squared plus V3 squared blah blah blah plus Vn squared. So basically take all the units that form this vector and then so are on this vector and use them, square them and then add them and then take the square root of that. That's the distance or I have to say the norm of this vector. We saw already the norm here is just another example what norm is and um on a specific two-dimensional vector when we have for instance that a vector is equal to three and four, which means for the first dimension, let's say on x-axis we have three and then on y-axis is equal to four, then the norm, so this is equal to we take the x value, so three and then we square it, so V you can see here this is the case when n is equal to 2, this is simply equal to square root for V1 squared + V2 squared and as V1 is equal to 3, so this is our maybe I can make this just V1 and this is my V2, then the norm or the Euclidean distance for this vector, so this thing is equal to V1 squared + V2 squared which is equal to 3 squared + 4 squared and this value is square root of 25 and it's equal to 5.

Let's now see the difference between Euclidean distance and the norm. So you could see here the norm here we have just one vector like here and this norm it has just two corresponding values into two-dimensional space. You see here we have just three and then four, so this is V1 and V2. When it comes to the Euclidean distance, this is kind of the generalization of this idea of norm. So the Euclidean distance between two points A and B in Rn, so in the N-dimensional space, is the norm of the vector connecting A to B. So we see that the norm and the Euclidean distance are highly related to each other, only we are talking about the norm when it comes to one vector, but when we have this vector A and the vector B, this is simply the Euclidean distance. So for the Euclidean distance we know already this idea of distance how we can measure it and you can see that this comes very similar to what we see here notation and here we are saying well we have this vector and then it has this two coordinates in n is equal to 2 in two-dimensional space. When it comes to the Euclidean distance, Euclidean distance helps you understand what is this distance between two points in an N-dimensional space. So the Euclidean distance between two points, let's say A and B in N-dimensional space is the norm of the vector connecting A to B. So for instance if we have a point A and we have a point B, we are connecting this and this is the vector connecting these two points, then the Euclidean distance is simply the norm of this vector. So this is the Euclidean distance. So we can see that norm and the distance they are highly related to each other. In the Euclidean distance we are using this idea of norm and specifically the norm two as I mentioned before. So here you can see that the definition of Euclidean distance, so the distance between A and B, the two points is equal to square root of A1 - B1 squared + A and then here we have basically A2 - B2 squared and then plus A3 - B3 squared, those are things that we cover as part of this dot dot dot and then plus after the last point when we have An - Bn squared. So here what we mean basically is that if we have two points here is A and here's B and this our vector and we know all these different points, so A1, B1, A2, B2, A3, B3 blah blah blah and then here An, Bn, we know all this points lie in here in this distance, then we are taking them and using them to calculate the Euclidean distance. So here for instance if we have um point A and B, so in this example let's do a quick one specific example when we have a point A which has coordinates one and two, so this is basically A1, A2 and then point B with uh points in it like B1, B2, you you can notice that the dAB, so the distance or the Euclidean distance of these two points which is equal to the norm of this um vector, but here this is A and this is B and this is this vector, this is equal to square root of 4 - 1, so it takes the B1, so this is B1 and this is A1, takes the square and then says plus B2 - A2 squared, takes the square root of that and says this equal to 5. Now you might be wondering but hey why do we do then instead of 1 - B1 squared we do B1 - A1 squared and the answer to this question lies in the um uh properties that we learn as part of pre-algebra because it doesn't matter when we take uh A1 - B1 squared or B1 - A1 squared because this squared ensures that it doesn't matter which one we take first and subtract the other. Now the proof of that is outside of the scope of this um course, this is part...

Of prealgebra, but I just wanted to put this out there to ensure that, uh, you are, uh, seeing what we are seeing here. Because here it says A1 minus B1, but in this example, we are taking instead depth, uh, B1, and we are subtracting A1. This is a common thing that we do in, um, pre-algebra and just in general, uh, in different, um, Alin dist or distance-related cases. So I just wanted to put this here to ensure that, uh, later on this is something that can be clear, um, from the first view.

Why this is important: this idea of norms and Euclidean distance, beside being used in machine learning, and why is it used? So norms, they provide a way to measure the size or the length of a vector in vector spaces. Which means that when we want to measure a distance, a similarity, a relationship between, for instance, vectors, then it becomes much easier to use this idea. And Euclidean distance is not only used in regularization techniques like L2 regularization or retrogression, but it's also used in other machine learning or deep learning algorithms as a way to measure the distance or the relationship or the similarity between two different entities. Those can be variables; those can be two people that we want to compare in our algorithm or two entities, uh, um, for instance, the, um, norms or the distance. They are also used as part of the K-means algorithm, something that you might have heard. And if you follow later on the machine learning and the clustering section of machine learning, you will see that Euclidean distance is used as part of K-means algorithm that aims to cluster observations into different groups. So this is also yet another highly applicable, uh, topic that you must know in order to understand different linear algebra topics, but also machine learning topics.

Why this prerequisites matter and why I mentioned those: understanding this concept is very crucial. They underpin this geometric interpretation; linear algebra; they will help you to better understand these concepts and not just to memorize them, but really understand. And later on, when you go into your machine learning and AI journey and your data science journey, seeing these concepts will help you to better understand those different algorithms; these optimization techniques; what we mean when we say we want our optimization algorithm to move towards a local minimum, a global minimum. This idea of movement, this idea of vectors; later on, we, you will also understand these different concepts in deep learning, how these models work, how the neural networks work. And those are essential concepts that you need for solving different systems of linear equations, a core part of this course. They also help in visualizing vector spaces, which are critical to understand this concept of linear algebra, the applications of linear algebra when it comes to the real-world applications. So those are things that you can definitely master by following some of our other courses, but for this course, I assume that you are already familiar with these concepts, right?

So now we are ready to actually begin. And with these prerequisites in mind, you are prepared to start your linear algebra journey. We are going to learn everything in the most efficient way, in such a way that you will learn the theory; you are going to see many examples; we are going to learn everything in detail, but at the same time, you're going to learn the must-know concepts. And I'm not going to overwhelm you with the most difficult concepts that you will not be seeing in your career. I'm going to give you the bare minimum when it comes to really knowing and the must-know for linear algebra such that you will be ready to apply linear algebra in your professional journey. Whether you want to get into machine learning, deep learning, artificial intelligence, data science, knowing these different concepts in linear algebra, you will be a pro in your field. Going to give you everything that you need: the theory, examples, implementations, everything in detail, but at the same time, you will be doing that in the most efficient and time-saving way. So without further ado, let's get started.

Hi there! So let's get started with our first module, which is foundations of vectors. In this module, we are going to talk about fundamentals of linear algebra, vectors. We are going to make a differentiation between scalars and vectors; we are going to define them. So first, we will learn the theory, then we will implement them into practice by plotting them, by looking into different examples. Then we will look into this representation of vectors by looking into the magnitude and the direction of it and the representation of them just in general. We are going to plot them in our coordinate system. Then we are going to see the common notation of vectors and indexing of them. Vectors are super important when it comes to linear algebra and application of it, and, uh, they matter not only in mathematics but beyond. So, uh, vectors help us in many ways, from figuring out how objects move to solving math problems in science and just in general in technology, including in data science, machine learning, artificial intelligence, etc. They are a super useful tool. So, uh, let's start our journey with looking into scalars.

So scalars, they are just numbers. And by definition, a scalar is a single numeric volume, often representing magnitude or quantity. For example, uh, scalars can be describing, um, the temperature outside, for instance, the temperature of, um, 22° can be represented by a scalar, or a height of a person can be represented; it's a scalar. So let's assume we have a scalar that we will define by a letter s; it's just a variable. This scalar is then equal to 22, for instance, and we are measuring it in degrees. So it means that, uh, if this s measures a room temperature, then and the scalar s, which is equal to 22°, it represents the room temperature. It can be, for instance, 18° or 9° if it's very cold; it just measures a single volume; it represents just a single number, or it can be, for instance, 17,100, 2.22. So all this, they are just scalars; they represent a single numeric volume; they often represent the magnitude or a quantity. Very soon we will see that scalars, they are a value that represents the magnitude of a vector. So, uh, now when we are clear on this very basic concept of scalars, let's actually move to this idea of vectors.

So by definition, a vector is an ordered array of numbers which can represent both magnitude and direction in space. So, uh, vectors, they are a bit more; they represent a bit more than scalars; they are numbers that also show direction, like a car speeding down the highway or a bow, uh, being thrown, for instance. Uh, when it comes to our previous example where we were using this, uh, uh, room temperature as a way to, uh, think about the scalar, a scalar, for instance, a scalar that we just saw was this room temperature, room temperature which was 22°. When it comes to the vector, a vector is different. For a vector, for instance, we can have an example when a bird, for instance, a bird, it flies at 10 kilometers per hour, and I also add here another information which will make this as a vector, which is that it flies south. So here, as you can see, what I'm doing is that I'm not just—oh, let me actually remove this part to make it easier to understand. Okay, so, uh, in this example, let me write it down that the example: a bird flies south at 10 kilometers per hour. So you can see that I'm not just adding the scalar, which is in this case the magnitude; we will see very soon the formal definition of it. So I'm writing down the speed; I'm defining the speed, but also the direction. So I'm saying I know that the bird is flying south; that's the direction, and I know also the speed of it, which is the magnitude. So 10 kilometers per hour. So here in the vector, I have much more information than in the scalar, because in the scalar I just got temperature, room temperature, single value, but in case of a vector, I not only have a magnitude or speed like 10 kilometers per hour, but I have extra information, which is the direction of it, for instance, flying to the south.

Let's now look into some real examples and plotting them to make more sense out of this idea of a vector and what is this magnitude, what is the direction. So let's assume we have a 2D plane, so we have an x-axis, we have a y-axis here, like usual, we have our zero center, and we want to plot a simple vector. So, uh, usually the way we represent a vector in tutorials or just writing down is by writing the name of the vector; this can be just a random name; let's assume that it's a v, letter V, and then on the top we are always adding this arrow. So this arrow, it says and it tells the person who is reading that we are dealing with the vector; the arrow on the top is that reference. So let's assume this, uh, vector v, it starts from the center of our coordinate system and it goes to this point. So let's say in here, this is our vector V. So let's assume that this point in here is equal to 4, which means that the x-coordinate is 4 and the y-coordinate is 0, as the, um, uh, arrow, it just—as the point in here, it has a y-value of 0. So you can see that it goes straight from 0 to this one, to this point. Okay. So what tells this vector, uh, to us is that we have a value that describes the length of the vector, so it goes from 0 to 4, which means that the length is equal to unit 4, so it's equal to 4, um, and we have just learned and we were just talking about that the magnitude is the length in this case. So the length describes the magnitude in this case, so this means that the magnitude of this vector is equal to 4. And then, um, what else we can see here? We can see the direction of the vector, which means that the direction is also something that we can see here; this is the direction of the vector, so this going straight from this point to this point in a horizontal way. So independent whether I plot this vector from 0 to 4 in here or in here or in here or in here or in here, in all cases, as long as the length is this, I'm dealing with the same vector, because I am basically in this entire R2 space; I have exactly the same vector; all I care is about the magnitude and the direction. Where will this vector start and where will it end? I am not interested; I'm interested that the—that the magnitude, in this case the length, is equal to the direction of the vector.

Let's now look into another example where we go a bit more difficult on our coordinates and on our vector. We already saw that we had this vector where we went—let me change the color—so this was our vector v, and it went from 0er till 4. So this point, to be more specific, is—so this vector, it goes—the vector v, it goes from 0, 0 to 4, 0. So the coordinate x was 4 and the y was 0. Now let's plot another one, um, where the direction is no longer horizontal. For this vector, let's call it vector w, and for this vector w, we will again start with 0, so we will start again in here, but this time we will go like this. So let's say we go all the way to this point. So this point has a value for an x-axis of 3 and for a y-axis it has a value of 4, which means it goes from this point to this point, and this is the direction of our vector v. So it goes to 3, 4, because this point is 3, 0 and this point is 0, 4. So the x-axis is 0 x-coordinate and y-coordinate is 4. So now you can see that the direction of this vector is like this, while the direction of the vector v was like this. And like in case of vector v, I again no longer care about where exactly my vector starts and ends, but all I care is about its magnitude, so the length and the direction. So for instance, I can have the same vector in here, the same vector in here, as long as the length, the magnitude is the same and the direction, I am dealing with the same vector; that's all I care. So the magnitude and the direction is all that you care about. All right. So now about the length, um, that's, uh, something that you can see very easily from this specific example because by using the Pythagorean theorem or Pythagorean theorem, we can see very quickly that as the length of this side of our right angle 30°, so right triangle, we can see that this side is 3, this side is 4, which means that this side is 5, because 4 squared + 3 squared, then we take the square root of that, a square root of 25, and it's equal to 5. So the length or the magnitude of this vector v is simply equal to 5. All right, this was about this, uh, specific vectors. Let's now look into the, uh, common representation of the vectors. So we always use the magnitude as well as the direction, you know, to represent the vectors, and they commonly are represented by two different, uh, ways. Let's now look into the first way that the vectors can be represented, and then we will move on to the next one.

So when it comes to the vector v, so we saw that vector v was moving from 0, 0 till, uh, to the point of 4, 0. So we can represent the vector v by (4, 0). When it comes to the vector w, we can represent that, uh, vector—so vector w, we can again do the parenthesis, and we can say that is equal to (3, 4). So by using the coordinates from the coordinate system, we can then represent our, uh, vectors. So this is just one way of representing a vector. Another way of representing these vectors is by using these square braces, given that we are in a two-dimensional space. First, we will mention here the 4, then we will mention the 0 in here, 2. So we can say [3, 4]. This is yet another way of representing the vectors in a two-dimensional space. So if we were to have a three-dimensional space, so let me actually show it on a new page. So if we were to—if we were—to have vectors in three-dimensional space, so we are dealing with R3, so we have points that can be described by x, y, and z, so coordinate space like this, so x, and then y, and then z, then every point—so let's say we have this vector—then we had to represent it by a value, let's say x1, y1, and z1, or—better, let me actually use different letters, a, b, and c, and this would be my vector v, and I could also represent this vector v is the same. So vector v can be represented as (a, b, c). So what thing that you can notice is that unlike the R2, now I have three different entries, what we are also referring as rows, and we just got one column. So, um, we can, uh, often represent and usually that's a common way of representing vectors by using these, um, columns. Columns help us to—to represent our vectors, and you can see very clearly then when it comes to the two-dimensional space, so when we have R2, so then our vectors have just two rows, so [3, 4], [4, 0], like in here. When it comes to a three-dimensional space, we have three entries, and so on. So the same holds, of course, also for, for instance, R5. Then for R5, our vectors, so coordinate space can be, for instance, x, y, z, and then gamma, and then let's say delta, and then the coordinates, uh, of a vector in that space can be v, and then arrow is equals to, and then we would have, let's say (a, b, c, d, e). You get the idea. So depending on the space, the coordinate space and the dimension of that space, then the corresponding vectors can be represented accordingly.

So the vectors are quantities that have both magnitude and direction, as we just saw, distinguishing them from scalars which only have magnitude. So we saw that the scalars got only magnitude, while in case of vectors we saw both for the vector v and for the vector w; we didn't—we didn't only have the magnitude, so the length of the vector, but also the corresponding direction. So, uh, when it comes to the, um, vectors, so this is exactly what we just saw in our example: a vector in a two-dimensional space, so in R2, uh, can be represented by using these square braces and the corresponding entries (x, y), where x is basically the x-coordinate in our coordinate system. So in our x and y system, whenever you have this x and y coordinate, then, uh, this x-coordinate will then describe your magnitude, and the y-coordinate will then describe your second entry that you need to put when representing your vectors. So here the x and y indicate the movement in the horizontal and in the vertical dimensions respectively. So for x's, it's always the x-coordinate, so how far you move towards the horizontal direction, in here, in here, or independent in here. So always take the x-coordinate; that is the value that you need to put first, and then the y needs to be put in here.

So indexing in vectors: when it comes to the, um, indexing, the standard mathematical notation, uh, indices in the n vectors goes from i = 1 to i = n. So the, um, notation here can be a bit ambiguous. So ai, uh, could mean the i-th element of ai, the a vector, or the i-th vector in a collection. So let's start with a simple one and then move on to this next part. So what this means and what this means we will look into now. So, um, usually, uh, when we have a, um, n-dimensional space, we are having a hard time visualizing it; therefore, we use this two-dimensional space or maximum three-dimensional space in order to get an understanding of what these vectors are. So we just saw examples of them, uh, when, uh, creating our vectors in, um, v and v, uh, and w in, uh, R2 and also in R3, but we can have similar vectors also in R4, in R5, or all the way down to Rn, where n can be 100, 200, 500, any number as large as you want. The thing is is that visualizing R4, R5, Rn is very hard, but we can still benefit from great properties of the vectors, matrices, and in general linear algebra in order to describe different things that have more than three dimensions. Therefore, we have this a bit more ambiguous notation where we use Rn, and this n can be any real number, and it can be all the way to infinity, so a very large number. And, uh, let's say we have a vector in this Rn, then this vector is usually described by using similar square, uh, brackets like before, only with, uh, more entries. So like before, we got just one column, so that's something that we didn't, uh, change, but here we have instead of just two entries or three entries like in the two-dimensional or three-dimensional spaces, now we have a1, a2, a3, all the way down to an-1 and an. So we got in total n elements in our column, and this describes our single vector. So this vector in an n-dimensional space, this we can call also a—. So one thing that we just saw is that it was saying in our definition and notation that, uh, we might also be dealing with the i-th vector in a collection, which means that sometimes you will see this. While here the a1, a2, they are vectors themselves. So here we saw that these are just entries, so a1 is a number, a2 is a number, a3 is a number, an is just a number, but it's also possible, uh, when you have a much more difficult and complicated case that you got an A—let's write it down with a capital letter A, which is equal to a1—or let's actually remove this—so we got let's say a1, a2, a3, all the way down to an-1 and an, where you can already see what is going on. So instead of having just a number as an entry, instead we have vectors in here. So our first element is actually a vector, our second element is actually a vector, so a2 arrow, a3 arrow, all the way down to an arrow. So while here this can be, for instance, some numbers, let's say 1, 1, 1, all the way down to 1, 1, here we have a vector, a vector, another vector, and all the way down here yet another vector, where, for instance—let me remove this part—where, for instance,

A1 error is actually equal to A1, 1, A1, 2, a13, all the way down to A1, n. One thing that you will notice here is that, unlike in here, here I got double indices, so I got here a one1 and then A1, 2, and then a13 all the way to a 1, n. So the first index it doesn't change, as I have here a one, so I'm writing down the index corresponding to this vector, but the second index it changes per entry, indicating which element specifically in the vector I'm talking about. So from the first index you can identify the vector that I'm referring to, which is A1, and from the second index you can see the corresponding um entry, or the value that that um element is positioned in this vector. So you can see that this value is, for instance, in the um vector one, so A1 to be more specific, but then it is in the first position; this is in the second position; in the third position; all the way down to the end position. So this is something that is really important to understand well, because this notation it's going to appear time and time again across various applications of matrices and vectors. So it is really important to understand well; therefore, I want to go one more time through this to make sure that we are clear on what this indices represent.

So whenever we have an index uh an a vector that we want to uh represent and it's um it has just um it is just a vector, which means that it's not a nested vector, vector in vector, then um we can define it by, let's say, a and then on the top an array, and it's equal to, and here we can have A1, A2, all the way down to a n. So you can see what we are also referring as dimension of this vector is equal to n by 1. So I got n entries and just one column, so n by one, which means that this already gives me indication that most likely this A1 is a number, this A2 is a number, this a n is a number. So let's say this equal to 1, 2, uh 3, blah blah blah, and then here I have let's say 100. But if I'm dealing with a nested vector, later we will see that this can be represented by a matrix, then um I can also define this by capital letter A, which is a common way to refer to either matrices or nested vectors, and then this is equal to A1 → A2 → A3 → This already sends a message to the reader that we are dealing with no longer um constants within uh vector, but rather vectors in a vector, and uh what can we see here is that the dimension of this nested vector, or which we can also refer to as a matrix here, the number of rows, so the number of entries, this elements, we can see it's equal to n, but then this time the number of values that form these vectors is no longer one, because we're are not dealing with just a constant, this is not some constant, but rather this is yet another vector. So let's assume this vector has a length of M, so let's say this has a length of M, then the dimension of this matrix a is equal to M, so something that we will see also when talking about matrices.

So let me actually clarify this bit more for better understanding. Let's say we look into one of those um one uh one other example of an entry. So let's say we look into this specific vector, which is in the the third uh vector within this vector capital A, so this A3 vector. So one thing to see here already is that I assumed that these vectors they got M elements, and keep in mind that all these vectors they should be of the same size, so it means that I already know that this specific vector A3 has M elements, elements, so M elements. So I'm representing this uh each vector from here, I'm taking this out from this entire uh nested a vector, and I just want to represent this, and now unlike this elements that got an arrow on the top, this time I will have uh constants forming the A3 vector, so I no longer have vectors, but I have elements in it. So in here I will have a, a, let me actually write down all the A's, but to refer and to make sure that I recognize that I'm dealing with the third a vector, so A3 → here I will put three, three, all the way here, three, so they all come from the same third a, three vector, but then their positions is different, because this is let's say uh one, two, and then all the way down to nth position. So this indices help us to keep track what are the um position that this values are taking part in the vector A3 → This might seem bit complicated at the moment, but once we move on onto bit more complex material like uh matrices it will make much more sense. This is bit of an extra, I just wanted to showcase this, but this is what uh is at its core and what you need to uh understand at the moment to understand this concept of vectors.

So you need to know that vectors can be represented by this arrow on the top, so let's say vector a, and it has let's say n elements, then you can write the square brackets, and then you will need to mention A1, A2, all the way to a n, which means that you have n different entries describing your vector. So you have A1, which is the first element in your vector, A2, the second element, all the way to a n, which is the nth element. But here you can see, for instance, so if I had here a three that uh A1 is simply equal to 1, A2 is equal to 2, A3 is equal to 3, all the way to a n is equal to 100. So these numbers I'm basically taking and I'm representing them, I'm putting them in here within square braces in order to get a representation of my vector. So my vector a has all these different entries and different entries, and it starts with one and it ends with 100. This is a vector, and then when it comes to the vectors within vectors, here we need to be bit more careful, cuz here we not just have uh constant values forming a vector, but we have vectors that form yet another vectors. So our vector a, our nested vector a, which we uh later will refer as matrix a, has actually entries that also are vectors. So we have A1 vector, A2 vector, A3 vector, they are not just constants, but on their own they are vectors. So here, for instance, we have defined also an example of it, we have said let's look into this third specific vector that is part of a, which is a three uh vector, and uh that one has M different elements. Here we have then the index referring to the which vector from the nested vector a it is, which is the third one, because we have taken it from here, but then on its own this vector has different members, and different members to be more specific, therefore we have also an index to keep track of the position of this value, one to up to M, and this can be yet another uh this time it can contain some elements, an example of which is, for instance, 0, 1, 2, all the way to let's say 500, and this can be different numbers, it doesn't need to be ordered, it doesn't need to have a specific pattern, they can be just random numbers describing this A3 vector.

So hopefully this makes sense; if it doesn't, don't worry, because we are going to see this time and time again. I just wanted to give you brief of an intro such that you can uh remember this when we come uh back to bit more uh complex topics like uh indexing in matrices. So now let's talk about about special vectors and operations. Here we are going to talk about zero vectors, unit vectors, the concept of sparsity in vectors, as well as vectors in higher dimensions, like we just saw about this n dimensional space. We will also talk about different operations we can apply when it comes to vectors, like uh addition, subtraction, and then later on in the next module we will also talk about multiplication. We will also be looking into the properties of vector addition after we looked into some detailed examples when it comes to operations on vectors. All right, so let's start with the zero vectors and unit vectors. When it comes to the zero vectors, you can see here already that um the zero and arrow on the top it basically refers to the vector like we saw before, only with the difference that all its members are zero. So you can see here that we have zero and then an arrow and then underneath here we have some number three, and then this is described by this common representation with the square braces and then three different members, 0, 0, 0, so all zero, and then it says in R3. Okay, so why are we doing this? Well, uh when it comes to uh different linear algebra operation, sometimes we just need to add zero uh vectors, or we just want to create zero vectors, it's just easier here to work with; we want to uh just create an empty uh vector, we want, we know the length, but we want to keep it empty, such later on we can add something on the top, or knowing that when we add a zero on a number the number stays the same, we can make use of this property to uh do different um uh tricks when it comes to programming in Python, in R, or in C++ etc. So therefore this idea of zero vector can become very handy. Now one thing that you need to notice here is that we are not just writing down this zero to emphasize we are dealing with a vector, but like uh before we have this arrow on the top emphasizing that we are dealing with a vector, then what we are doing is that we are also adding the dimension of this vector, so in what dimension, in what space are we um uh creating this zero vector, that this vector is located, is it in R2, in Rn, in R3? In this specific case you can see that in this example the uh index that we got here is three, which basically indicates we are dealing with a zero vector in three-dimensional space, so in the R3. In general we would just note this by n, keeping the uh notation general, which means that we are dealing with 0, 0, all the way down to 0, so it has n by one dimension in r n. All right, so this is about zero vectors; it is just a way to uh make our programming life easier, also to use it in different uh algorithms when it comes to bit more advanced algebra.

The next type of special vectors that we will look into is this unit vectors, so vectors with a single element equal to one and all the other zero, denoted as Ei for the E unit vector in N dimensions are referred by unit vectors. So uh what we mean here when it comes to the unit vectors, uh if we have for instance E1, it means that we have a vector where the is in this case the first element is equal to 1. So you can see that E1 is equal to 1, 0, 0. So in the first element we got one and the remaining is zero, and this is really important that we are dealing with vectors that contain only elements of zeros and ones, and the only member that is equal to the only element in that vector that is equal to one is the ith element in the entire vector, all the remaining ones are zero. And you can see here that the dimension is no longer specified, but just the um index of the entry where the um uh the uh one is located. So let's look at another example in here, for instance, when it comes to the um uh unit vector, yet another unit vector is E2, which basically means that in the second element, so in the second place uh the uh vector contains one and all the other members are zero. So here you can see zero, here you can see zero, only in the second element we have one, and then in the E3 what we have here is that the third element is one and all the other ones are zero. So let's actually look into uh one um bigger vector uh in higher dimension to make it even more sense. So first I will define and assume that we are dealing with a vector in Rn, so in an N dimensional space, this gives me an idea that we are dealing with um so we are not dealing with nested vector, we are dealing with a simple n dimensional vector, so it has n rows and one column. So using the square braces I'm going to represent my vector, so I have all these different members, n members. So e, let's say it is E5, so what does this mean? It means that i is equal to 5, and this ith element, so the fifth element is equal to one, and all the other entries, the elements in this vector are zeros. So let's look into this: 0, 0, 0, 0, I'm approaching the fifth element in my vector, so it's this one, this is one, and the remaining all zeros. So this is a unit vector in an N dimensional space, and I'm defining it by E5 because my fifth element is equal to one. Now those are very handy when it comes to some other uh techniques in linear algebra and just in general, think about techniques like um uh row echelon form, solving linear equation, something that we will see as part of the next unit. So many things um we can do by using unit vectors. Unit vectors are super important, so you need to understand this concept uh very well such that later on you will understand uh more advanced concepts in linear algebra.

Now let's look into the topic of sparsity in vectors. So by definition, a sparse vector is characterized by having many of its entries as zero. So its sparsity pattern indicates the position of a non-zero entries. So uh what we are basically saying is that if we are dealing with a vector that contains too many zeros, we are dealing with a sparse vector. So uh this sparsity pattern indicates uh also the positions of a nonzero elements. So um if we have um unit vector, it means that we are already dealing with a sparse uh vector. This is a concept that is super important when it comes to linear algebra, but also in general data science, machine learning and AI, because having the sparsity in your vector it means that you don't have much of an information, usually a value is zero it means you don't know much about that specific value, and if you got just too many of zeros and too few numbers which do um provide information, it means that you are dealing with a vector that doesn't provide you much information, and there's always a problem when it comes to data science, machine learning and AI. So sparsity is something that you need to be aware of, you need to know how to recognize it, and you also need to know whether there's a problem in your specific case or not. So let's look into an example. Let's say we are dealing with this vector X that has five different elements. So X is a vector coming from um five dimensional space. So we have for instance an element of three in the first entry, then we have zero in the second and third uh entries, then we have an entry um four, which coincidentally also contains value four, and then the last element in our five dimensional that vector X is equal to zero. Now what do we see here? We see that the majority of elements of a vector X is equal to zero, because we got in total five elements, and then we got three of it actually uh being equal to zero, and only two of them containing information, like equal to three and four. So only two elements that are not zero, so nonzero elements, it means that 3 / 5, which is is basically 60%, 60% of all the entries in the vector X are equal to zero. So the 60%, it means that is above half, so above 50%, 60% of all the information in this vector um the majority is simply equal to zero. This type of vectors we are uh calling sparse vectors, and sparsity is really important concept uh that we need to keep in mind later on.

So uh while we can visualize vectors in two and three dimensions in linear algebra, like we just saw in case of this n dimensional vectors, visualizing uh the this type of higher dimensional vectors becomes very difficult. So uh this mathematical flexibility uh to work with uh this type of uh information, so when we can represent uh information with many entries, we can represent it by vectors which we can actually not visualize, becomes very handy for complex data structures, for different simulations in physics and much more. So uh we just saw in couple of examples uh how we can represent vectors in a high dimensional space using this square braces and this common vector notation representation. We saw that in an N dimensional space we could uh very easily represent this uh very large matrix or vectors uh by just um using this vector representation. For instance, if we got a vector that had n different entries, where n is for instance thousand, so let's say we have thousand, then uh we can represent uh this uh vector or this information by using common vector notation, so A1, A2, all the way to A1000. So of course we cannot visualize this, it just doesn't make sense, we can visualize two dimensional vectors, we can visualize three dimensional vectors, but we cannot uh visualize thousand dimensional vectors, so vector that comes from uh R1000, but what we can do is still make use of this very useful information in order to uh do different operations, when and later on we will see that uh this property, as specifically this part of linear algebra, it helps us to work with vectors in any number of dimensions, whether thousands, millions, billions. This mathematical flexibility is super important for more complex data structures, uh for matrix multiplications, when for instance we are doing different uh algorithms, including how we can represent a very large matrices, very large feature spaces, all this different information we can represent just by making use of vectors coming from this specific uh part of linear algebra.

Let's now finish off this module by looking into some applications of vectors. So one common application of making use of vectors is uh when we are performing different operations while having words and we want to count those words. So this is a super common application of vectors, and we can account this words and can even plot a histogram of how often each of these words appear in a document. So a vector of a length n, for instance, can represent the number of times each of these words in a dictionary of n words appears in a document. So uh just for the sake of simplicity, let's assume that we got um dictionary that contains only three words; of course, in the reality the um dictionary, what we also refer often as corpus, it contains much many uh much more many words, but for the simplicity we will assume that we just got three different words in our dictionary, so that's a total. Now let's assume that we got a document uh with these different words and we want to count how many times each of those words that we got in dictionary actually appear in our document. So um if our document is described by this vector, so it contains an entry of 25, 2, and 0, it means that in our document we got 25 word one in our from our dictionary, so in the position one, two times word two, and zero times word three. So basically we have a predetermined set of words in our dictionary, in this case three words, word one, word two and word three, and they have a specific index, specific position in our vector, and when we are putting these values in here, then the machine or the uh computer, the program will understand that if we have 25 in the first position, then the word one in the dictionary appeared 25 times in our document, whereas the second word appeared only two times, and the last word, word three, didn't appear at all, so zero times, times in the entire document. So let's look into a practical example actually to make even more sense. So um this is by the way a common practice to count different variations of a word; they are common application in n-grams, large language models, transformers, they are just a cornerstone of many language models when we want to count the words in the document to understand how often the word appears, because this gives us an idea what this document is about, knowing how many times the same word appears in that uh document.

It gives us an indication of the topic of, uh, the document. Also, we can make use of it to do sentiment analysis to understand what this document is about, not only in terms of the topic, but also is it a positive, is it a neutral, or a negative, uh, document, so to say. So, uh, for instance, if we got, uh, the following words, uh, that correspond to our dictionary, and in our dictionary we got just, um, let's say six different words, then what we can do is that we can say 3, 2, 1, let's say 0, 4, 2. And the corresponding words are: word, row, number, horse, eel, and then document. What this means is that we have a text, what we refer to as a document, that contains three times the word "word," that contains two times the word "row," contains one time the word "number," zero times the word "horse," and four times the word "eel," and two times the word "document." So, uh, this is basically a common way of representing the, uh, frequency of the words in the document.

Let me actually give you, uh, another example. And in here, I want to emphasize another thing: the concept of stop words. So, uh, let's say I make this 10, and then here I say there are three times the word "I," two times the word "reading," two times the word "library," four times the word "book," zero times the word "shower," and 10 times the word "uh." So, uh, you can see that in here we are dealing with a document that contains 10 times the word "uh," which is what we refer to as a stop word. So those are things that actually don't give us too much information about what the document is about because, uh, it's just used everywhere, but it is appearing too often. So you can see 10 times—the most frequently appearing word. This is what we refer to as a top word. And then another thing that we can observe—the second thing we can observe—is that we are dealing most likely with a document that describes library reading, uh, because you see words like "reading," you see words like "book," "library." But another word, "shower," that is totally unrelated to reading, book, or library, is appearing zero times. So even by looking at these counts, we can already get an idea of what the topic of this document is about.

So, uh, you can see already now from this very basic example, where I made too many assumptions regarding how small the, the, uh, dictionary should be, you can even see now how we can use these counts in our dictionary from our text in order to get an idea about the topic of the document or topic of the conversation. It can be a topic of the, uh, tweets if you have tweet data; it can be a topic, uh, regarding books if you have many book, um, uh, book texts; it can be, for instance, the topic of the review if you got, uh, product reviews from, uh, Amazon, for instance. Using these counts can help you to get a topic—regarding topics—from that text. Then you can also use it to remove the stop words because usually the stop words are the most frequently appearing words. It can also give you an idea about the sentiment. For instance, here we are dealing with neutral sentiment; it's not positive, it's not negative; it's just reading a book in a library—that kind of topic. So all this can be super helpful when it comes to natural language processing, that's a field where this, uh, text processing—text mining—and then using that for modeling purposes is what, uh, what plays a central role. It also plays a super important role in the large language models, in the Transformer models, and, uh, in the simple matters like, uh, bag of words or, uh, in the, uh, tf-idf. All this is based on this idea of counting words and how we can use this information. And you can see how vectors come into play in these different applications of linear algebra in data science, natural language processing, artificial intelligence, and machine learning. So they are super important.

Another application of vectors can be representing customer purchases, for example, an n-vector p. So, let's say p can record a customer's purchases over time, with pi being the quantity or dollar value of an item i. Now, what does this mean? So, let's say we have vector p that represents the customer purchases, and we are dealing with a single customer, and we are just saving over time that information—how many times this customer has made purchases over time. So the quantity is in, um, dollars—the dollar value of item i purchased. So we are basically keeping track of what is the value of the item i that the customer has purchased. So what we can do is we can assume that in here, actually, it already makes that assumption; it says n-vector, which means that the number of rows or number of, um, items that the customer purchases is n. Now, what the, um, the problem says that it represents is that in each entry—and here we have in total n entries—we got a dollar value of item i, which means that here, if I have p1, p2, all the way to pn, and here somewhere in the middle I got pi in the i-th position, it means pi represents the value of item i. So, for example, if I'm dealing with a customer that buys, um, let's say, uh, courses, and the first item that the customer is buying is a mathematics course, so I'm writing "mathematics course," and this is the first course that they buy—i is, by the way, just a, um, way to refer to the i-th purchase. So let's say, um, here somewhere in the middle the, um, customer decides to buy a deep learning course—"deep learning course"—and then it continues by buying, uh, the customer continues buying courses, and the last course that the customer buys is, let's say, um, a career coaching course. Now, let's say the mathematics course costs, uh, around $1,000; let's say the, uh, deep learning course costs $3,000; and then let's say the career coaching service, which is usually one of the most applied and personalized ones, can cost all the way to $5,000. Now we see that in the i-th position—this is the i-th position; let me change the color, by the way—so let's say this is the i-th position, this is the first position, and this is the last position. So those are just indices. We can see that in the i-th position we got the 3,000, which means that the pi is equal to $3,000. So this indicates that in the i-th purchase the customer purchased a deep learning course, and the value of that item was equal to $3,000. All right. So now we are ready to go on to the next major topic, which is about vector addition and subtraction. So we are going to do some operations and apply these operations to vectors.

So let's first formally define this idea of vector addition. Two vectors of the same size are added by adding their corresponding elements; the result is a vector of the same size. So, uh, let's unpack this. It says two vectors of the same size are added by their corresponding elements. So here it refers to two different vectors, let's say vector v and vector w, and it says let's add them—what we refer to as vector addition. And it says for that what we need to do is to take all the elements of v and then all the elements of w, and using their corresponding elements—indices that help us to understand where those elements are located—we are using in order to add each element in the vector v to the element of the vector w in the same position. And do note that in the second part it says the result is a vector of the same size because we are adding two different vectors of the same size. Mentioning here it means if we add two different vectors that have the same size, we are going to end up with a vector that has the same size. Now, once I go into the examples, it will make much more sense. Let's quickly also look into this concept of subtraction. On its own, subtraction is very similar to this idea of addition. So if we have a subtraction, let's say we have vector v, we subtract vector w, then we are doing basically, uh, what we just did in the addition, only instead of, uh, doing add, we are doing subtract. So again, we are just, uh, we are just subtracting from vector v vector w; they have the same size, so we end up having the result, which is a vector of the same size. Only one thing that you can see is that this can be also rewritten as v vector plus, and then minus w. So we basically can represent subtraction, um, on its own as a way of adding, only we take the negative—the, um, opposite directed vector. So this will make even much more sense, uh, once we go on to the examples. So let's look into our first operation example where we are adding two different vectors. This is a basic example; we got just two-dimensional vectors. We got vector a that has entries 2, 3, and vector b that has entries 1, 4. And what we are doing is that we are adding vector a to vector b. We just learned that we need to have the same size of vectors. So you can see that vector a has a dimension 2 x 1; vector b has a dimension of 2 x 1. So their sizes are the same; both have two entries, only two elements. And at the same time, we just learned that what we need to do is to take their corresponding elements and add them to each other. Now what does this mean? It means that we take from a the first element, 2, and then we take the first element of the second vector, which is b, so we take the 2 from here and 1 from here—the first element of a and the first element of b—and then we are adding them to each other: 2 + 1 = 3. And then the same holds for the second entry. So 3, which is the second element of vector a, and then 4, which is the second element of vector b, we are saying 3 + 4 = 7. So let me write it down even in a simpler manner such that it will make much more sense. So vector a has elements 2, 3; in the first element we got 2; in the second element we got 3. So a. Then we want to add b, which has in the first element an element equal to 1, and the second element is equal to 4. This means that if we want to add these vectors, (2, 3) + (1, 4), this is equal to—we need to take 2, we need to add 1—so this element and this element—and then we need to take 3, we need to add to 4—so this one and this one—which is equal to: 2 + 1 = 3; 3 + 4 = 7. So we got a vector (3, 7). Do note that this vector—the result vector—contains again two elements and just one column, so 2 x 1. So you notice that the size is the same for this result vector.

Now let's actually generalize this concept before moving on to the next example. If we got, let's say, vector a that contains n elements, a1, a2, all the way down to an, and it is from n-dimensional space, and we got vector b that also has n elements—remember that they both need to have the same size—so b1, b2, all the way to bn, so they come also—b comes also from n-dimensional space—so then when we add a to b, this is equal to a1, a2, all the way to an, plus b1, b2, all the way to bn. So n x 1, n x 1—the sizes—this is equal to—let me actually use this color to make it even more visible—so for the first entry for my result vector, I will get a1 + b1, then a2 + b2, so all the way down to the nth element, which is an + —let me use a different color—a1, b1, b2, bn. So you can notice now in general terms what we are doing here. So we are taking the a1 coming from the vector a; we are adding in the same, uh, position the value that comes from vector b, which is b1; we are saying take the a1 + b1; this is the, uh, first element, so the position stays the same, and then in the result vector. So we take all the corresponding values that are in the same position in the corresponding vector—first from vector a and then vector b—we are adding them, and this forms our new vector. And this new vector will again have a size n x 1. So you can see that the sizes of the two vectors are the same; both have n elements; and then we are using their corresponding elements to add them to each other element-wise, and then we are getting the result that has the same size, so n x 1. So this is a more general description of how you can add two vectors. Let's now look into this specific example. So we have a vector with the entry (0, 7, 3); so this comes from R3—you can see—so three-dimensional vectors; the second vector is (1, 2, 0); and then the final result is (1, 9, 3). So how we got this: we took 0, we added 1; 7, we added 2; and then 3, we added 0. So you can see all these elements element-wise, and then this is equal to: 0 + 1 = 1; 7 + 2 = 9; and then 3 + 0 = 3—exactly what we got here. So again, the same sizes, and the result is from the same size—quite straightforward.

Now, when it comes to vector subtraction, what are we doing there? So what are we doing here? So we are doing kind of a very similar thing; we are taking this element 1, we are subtracting the other one in this first element; then we are taking the 9 in the second position and subtracting this again from the second position of the second vector, and we are putting in here 1; and then 1 - 1 = 0; 9 - 1 = 8. So we get result vector (0, 8), like in here. And you can, you can see that the sizes stay the same; so also in this case. Let's write a more general, um, this idea of subtraction. If we got a vector a from R<sup>n</sup>—so n-dimensional space—and it can be represented by a1, a2, all the way down to an—so it has n elements, n x 1—and then we got b also from R<sup>n</sup>—so coming from the n-dimensional space, which means that it's got n elements—so b1, b2, all the way down to bn—again with the same size, n x 1—then a - b is simply equal to a—a1—let me actually use the same colors to make it easier to follow—so let me first draw my square braces, and then here I will use blue for a and then red for the, uh, color for the second vector, which is b; here I will use black—minus—then, given that the same size should be for the result vector, I already know that I expect n different elements for this; and then here I'm taking this first element that comes from vector a, I'm subtracting from this the first element that comes from vector b, so element-wise subtraction—b1—and I'm already getting the result for the first element in my result vector. So you can see a1 - b1. I'm taking this element and this element and subtracting them from each other to get a1 - b1, and then the same holds for all the other values, only coming from different elements from vector a, subtracting from this the corresponding values element-wise from the vector b—b2, b3, all the way to an—so you can see that in my result vector, a vector - b vector, in the first element I get a1 - b1, then a2 - b2, then a3 - b3 in the third element, all the way down to the nth element, which is equal to—oh, this should be b<sub>n</sub>—so, um, this already should make, uh, much more sense. Every time we take the element in the same position from one vector than the other, we subtract them from each other in order to get the corresponding element in the final vector. All right. So let's now, uh, before moving on to the properties, um, I wanted to show you, um, this only in a coordinate space. So what this means in terms of visualization in a coordinate space. So, uh, let's say we have a coordinate space; this is my y-axis, this is my x-axis; so this is x and y, and this is my center, so (0, 0). And what I'm doing here is simply I want to have vector a, let's say this is just, um, vector a—simple one—with the coordinates, um, let's say (4, -2), and I got vector b—let me use a different color—vector b that has coordinates, let's say (-4, 4). So let's actually visualize them. Let's first start with the vector a, uh, which has an x-value of 4—1, 2, 3, and 4—and the y-value -2, so this is my vector a. And let's now visualize the vector b, so (-4, 4), which means that—let me actually extend this—this is -4, so the x-coordinate is -4, so it should be here, and then the y-coordinate is 4—1, 2, 3, and 4—it's this one, which means that my vector b is this one. All right. So you can see now that the vector a is in here and the vector b is in here. Now what I want to do is to add these two vectors to each other. So what I want to do is to take this vector a and add to this the vector b, which is equal to 4 - 4 = 0, and then -2 + 4 = 2, so (0, 2)—it is (0, 2). So this is my result vector. So now when we are clear on how we can add vectors, how we can perform these different operations, and what it means in practice when it comes to looking at the vectors in a coordinate space and adding them or subtracting them, we are ready to look into the properties of vector additions. This is something that will definitely seem familiar to you, uh, from pre-algebra, where we are basically using all these properties that we already know that hold for, uh, numeric values—for the scalars—that are being transferred to this vector space. So we are going to talk about these four different properties that the vectors have. The first one is the commutative property, which says that if we add a vector a to vector b, then this is the same as adding a vector b to vector a. So basically the order of the vectors doesn't really matter when it comes down to adding them. So formally, a + b = b + a, for any vectors a and b of the same size. Then we have the associative property, which says a + (b + c) = (a + b) + c. We can write both as a + b + c. Now what does this mean? We know from pre-algebra that this parenthesis means first do this addition and then do the rest of the operations in here. It basically says if you add a to b first and then you add c, it's the same as first you add b to c and then on top of that you add the vector a. So then the third property is addition of zero vectors, which says if we add a zero vector to vector a, then this is equal to adding a vector 0 to a, and this is equal to vector a. So adding the zero vector has basically no impact on the vector whatsoever. Then the final property is subtracting a vector from itself, which means if we take the vector, we subtract the same vector from itself, so a - a, and we get a zero vector. So a - a = 0 vector, and this yields the zero vector.

Now let's look into each of those properties one by one and let's, uh, look into specific examples. In some cases we will prove this on the example that we have to make these concepts much more clear. So let's start with this commutative property of vector additions. So we want to see whether a + b = b + a. So let's say we have a vector a that has coordinates or magnitude and direction that is equal to (1, 2). Then we

Have a vector, uh, let's say B, that has a magnitude and direction of -2 and 3. So the first thing that we want to check is indeed whether A + B is equal to B + A. Therefore, let's first calculate this part, and then we will calculate this part that I will define by one and two, and we will see whether we are indeed having the same value, the same vector, or not. So let's see. So we have here A.

A + B, which is the first value that we want to calculate, A + B is equal to (1, 2) + (-2, 3). And we learned before that this is simply equal to take this value and then add this one: so 1 + (-2) and then 2 + 3. So this gives us a vector: 1 + (-2) = -1, and 2 + 3 = 5. So we get that A + B is equal to (-1, 5) this vector.

Now let's look at the second quantity: so vector B plus vector A. This is equal to (-2, 3) + (1, 2), and this is equal to -2 + 1 and then 3 + 2. This gives us -2 + 1 = -1, and 3 + 2 = 5. So we can already see from here that the quantity one is indeed equal to quantity two, which proves that indeed A + B is equal to B + A.

What this basically means is that adding two different vectors, the direction or the order is not important; whether you add A on top of B or B to A, it doesn't matter; the end result is the same. And actually, you can also see it if you, uh, combine this or if you do this in more general terms.

So let's say if we have a vector A, which is equal to, in an N-dimensional space, (A1, A2, up to An), so it has n x 1 dimension, and you have a vector B with the same size from the same R<sup>n</sup> dimension, and it has elements (B1, B2, up to Bn), and the dimension is equal to n x 1, then if we calculate first A + B, and this is equal to simply (A1 + B1, A2 + B2, up to An + Bn), and if you calculate the second amount, which is B + A, and this is equal to (B1 + A1, B2 + A2, up to Bn + An), you can see that A1 + B1 is equal to B1 + A1; simply from pre-algebra, you know that if those are all constants, for instance, 2 + 3 = 3 + 2. In the same way, A2 + B2 is equal to B2 + A2, and then here up to An + Bn is equal to Bn + An. What this means is that all these elements, they are basically the same, which means that we already have a proof. So we get this proof, and we can see that even for the general term, independent of what this vector A is, what this vector B is, that A + B is equal to B + A. This is exactly what we saw before in the first property, which is called the commutative property of the vectors, that A + B is equal to B + A.

Now let's move on to the other property, which is called the associative property of the vectors. Now what this property does and says is that (A + B) + C is equal to A + (B + C), and this is then equal to A + B + C. Now let's then see this specific property on an actual example. So what is basically says is that if we have this example where A is equal to—actually, I had this before; let me simply just remove this part. Let's then add our third vector, which is C, and let's call it—let's say it has a representation of (4, 5). Then the idea behind this property is that what we need to prove here is that (A + B) + C is equal to A + (B + C), and this is equal to A + B + C. Let's see actually whether this is indeed true for this specific case. Now this should come very intuitive, so I'm going to do it very quickly. So first we have this quantity, this one, then we have this one, and the third one. Let's do it very quickly.

(A + B) + C is equal to (1, 2) + and then we add C, so it is simply (4, 5), and then this is equal to—we saw before when doing this that we were getting (1 - 2, 2 + 3)—and then we add this (4, 5). This is simply equal to: 1 - 2 = -1, and 2 + 3 = 5 + (4, 5). Now, given that it doesn't really matter—no longer—that do we have parentheses or not, this basically means that this value is simply equal to -1 + 4, so here -1 + 4, here 5 + 5. So this is then equal to (3, 10). All right. Let's then now quickly do the second amount, which says: first add the vector B to vector C, and only then add the vector A on top. What this means is that we need to take (1, 2)—this is vector A—and we will only add this once we have added the (-2, 3), the vector B, plus the vector (4, 5). Okay, so we can see that we are just leaving this in here. Let's first add these two: -2 + 4, 3 + 5. So this gives us (1, 2) + (-2 + 4 = 2, and 3 + 5 = 8). So this gives us—let me remove this calculation—so this gives us 1 + 2 = 3, and then 2 + 8 = 10. Okay, great. So now we got already the quantity one being equal to quantity two. Let's check whether this is all equal to this one. It should already be something that you see now. Given that we know just from mathematics that parentheses doesn't really matter when it comes to the scalars, and adding two vectors is basically very close to this idea of the additive property of the additive property of the scalars, but just let's quickly do it to be 100% sure. So when we take this vector A to the B and to the C, we had all this; this is equal to (1, 2) added to (-2, 3) and then added this to (4, 5). Now what this is equal to—let me actually write this in a bit shorter way such that it can be all fit in in the small place—so (1, 2) + (-2, 3) + (4, 5). This is equal to basically 1 - 2 + 4, and then 2 + 3 + 5. Now what is this number? 1 - 2 + 4 is simply equal to 1 - 2 = -1, and then + 4 = 3. So the first element is 3. 2 + 3 + 5 = 5 + 5, which is equal to 10. Perfect. So now we get the confirmation that (A + B) + C = A + (B + C) = A + B + C.

So let's quickly also look into this addition of zero vector and the subtracting a vector from itself properties, and the detailed explanation of this or example of this I will leave it to you. So when it comes to this: A + 0 = 0 + A = A. So this property—let's say if A is equal to (2, 3), and then we are adding on this A + some zero vector, which basically means take (2, 3) and then add the same size of zero vector—you can see that this is the same as adding these zeros on these values. Now what do we get? We get that this is equal to 2 + 0 = 2, and then 3 + 0 = 3. There we go. So we already see very quickly that it doesn't really matter whether we add a zero vector to this original A vector or not; in all cases, it just adding a zero vector has no effect. And seeing from the commutative property that A + B = B + A, we already know that if A + 0 = A and is equal to this, then also 0 + A will be the same. And we can see indeed that we just saw that A + 0 is simply equal to A. So we basically have quickly proven all this. Now when it comes to the subtracting vector from itself, I think this is a very nice one just to see how we, uh, take the same vector and subtract from that value, and we get zero. And this is very similar to working with just real numbers; in the same way as 3 - 3 = 0. Also, when we have a vector consisting of the scalars, like A = (2, 3), in the same manner, if we take this A and we subtract it from itself, so A - A, then what we will get is (2, 3) - (2, 3), and this will give us 2 - 2 = 0, and then 3 - 3 = 0. So we get a vector zero, so zero vector.

So now when we are clear on how we can perform different operations on our vectors, and also we know the properties of adding and subtracting different vectors, we are ready to move on to a bit more advanced topics. So in this module, we are going to discuss this idea of scalar multiplication; we're going to look into the example how what happens and how we can do the vector multiplication with the scalar; then we are going to look into the span of vectors, what it means to have a span of vectors, what is the linear combination and the relationship between the span and linear combination and the unit vectors; then we are going to look into the application of scalar vector multiplication in audio scaling example; and then finally, we are going to finish off this module by looking into the length of a vector and a dot product, and we are going to go back to this idea of distance, understanding vector magnitude, and understanding vector length. So let's get started.

Now, before we look into this idea of span and linear combination, I quickly wanted to look into this idea of scalar multiplication and the specific definition of it. So formally, the scalar multiplication involves multiplying each component of a vector by a scalar value, effectively scaling the vector's magnitude. So what do I mean here? Let's say we have a vector, and I will write it in the general terms to keep everything general. So let's say we have a vector A—let me pick up my pen—A, and this vector A is from n-dimensional space, so it is from R<sup>n</sup>, and it can be represented by (A1, A2, up to An), and I have this magnitude of a vector, and now I want to scale this vector, for which I know the magnitude and the direction; I want to scale it with a scalar, and we learned before that the scalar is just a number. So scalar, in this case, I will be referring it to by C, so C will be my scalar, and this comes from R, which means that it's a real number. Let me actually use a different color to make it easier to follow. Okay, so my scalar will be with the color red, so C, and C comes from R. So what do I mean by scalar multiplication? I mean that I want to find what is this C times A. This is what we mean by scalar multiplying with a vector. Now what does this definition say? It says when we are multiplying a scalar with a vector, so the scalar multiplication, meaning multiplying a vector with the scalar, it involves multiplying each component of a vector by a scalar value. So if we translate it to this specific example, it means that this amount, so this amount is equal to taking C and multiplying it with each element of this vector, so each component of a vector. And what are the components of my vector? The A1, A2, up to the point of An. So all these components. So that means that the first element of this new vector, the scalar multiplication result, will be C * A1, C * A2, dot dot dot, so all this middle elements, and at the end again C times, and then An. And then in both cases, of course, the number of elements doesn't change, so the number of rows of my vector doesn't change; it's n. So here also n, and then the number of columns is the same, so it's just a column vector, so one column. So what we see here is that we go from A1 to C * A1, we go from A2 to C * A2, up to An transforms into C * An. So we see very easily that I keep all the elements from this vector; I take them in here, and instead what I'm doing is that I'm multiplying every element from this vector by the scalar C. So this is exactly what this definition says. And let's actually go ahead and do a hands-on example with some real numbers to have this method and to have this definition very clear in our mind, because we are going to make use of this fundamental operation, scalar multiplication, on and on in the upcoming lectures and just in general in your journey in any applied sciences.

This is an example of scalar multiplication. Here, here what we are doing is that we want to multiply this vector C—so in this case, the vector is defined by a letter C, and then on the top we can see the arrow indicating that this is the vector now—and here we refer the scalar by a letter K. We are saying we want to perform scalar multiplication, which means that we want to multiply the vector C by the scalar K. So how we can do that? So what we want is to multiply K by C, and we just learned that for that what we need to do—let me write this over—this equals -2 multiplied by (4, -3). This is my vector, so this is the K, and this is the C. This is equal to—so I take my scalar and I multiply it with each of the element of the C—so -2 * 4 and then -2 * -3. So -2 * 4 = -8, and then -2 * -3, so -- it goes away; it becomes a plus, and the 2 * 3 is 6. So my end result, the K * C, is equal to (-8, 6). This is my final result.

So let's quickly also do yet another example, and this one is a unique one because it's relating to this idea of you multiplying something with a zero, which is something that we also know from our high school that when we multiply a number, let's say seven, by zero, they're getting zero. And here in this example, the problem is: describe the effect of a scalar multiplication by zero on any vector v or r, which means what we are doing is that in this example is we want to know what is this result of any vector, let's say vector C. So we will use the same example C, only this time instead of multiplying it with the scalar k = -2, our scalar will be zero, which means that C = (4, -3), and then K is now equal to zero, and we want to find out what is this K * C. Let me actually write down the K with a different color: k = 0. So what we want to find out is K * and then C, and this is that equal to zero. So I'm taking the K, 0, times then I'm taking each of the elements of C, which is 4 and then -3, and I know that when multiplying the number with a zero it gives me zero, which means that I end up with 0 here. 0 * 4 = 0, 0 * -3 is also zero. So I end up with a zero vector. Now this gives me an idea already that I can make a general conclusion that independent of the type of vector that I have, independent what are these values in my C, if I have any vector C and I'm multiplying it with zero, then this will always give me a vector of zero, because all the members of this final vector will be just zeros. So if, for instance, the C comes from, let's say R<sup>n</sup>, so it has n different elements, it comes from n-dimensional space, then my final result of 0 * C, so this zero vector, this one, so zero, that this one will come also from R<sup>n</sup>. So you will be having a vector, so 0 * C will then be equal to (0, 0, blah blah blah blah 0), so n times zeros. So this is then the idea of multiplying, so scaling a vector with zero, and this is our example to.

All right, so let's now move on to our application of scalar vector multiplication, and then after this we will go back to this idea of linear combinations and spans. In this specific application, we have a scalar vector multiplication, and we are looking into the application of audio scaling. So the scalar vector multiplication in audio processing, this can change the volume, for instance, of an audio signal without altering its content. So you might have noticed that when, when you are listening to video, you can simply increase the volume of that video or decrease it, but you will notice that the content doesn't change; you are just increasing the volume or decreasing it. Even on the TV, when you are watching a show, you are increasing the voice or decreasing. Now what you're basically doing behind—and this is super interesting—is that behind the scenes what is happening is that there is simply audio that contains that show, and the audio of that show is being multiplied with a scalar, and that scalar is simply the volume scale. If you scale it in such a way that you want to decrease the volume, so the audio will then have lower volume, then you are simply multiplying your vector containing the audio information in such a way that those newer volume indications, they will be they will be containing lower numbers. Hope this makes sense. Let's look into the example; this will definitely clear this out. So let's assume we have a vector A that represents the audio signal, and we want to multiply vector A by scalar B to adjust the volume. So B is some sort of number; it can be—so B comes from R, so is a real number, while A is simply a vector. Given that it doesn't mention it here, I'm assuming that A comes from R<sup>n</sup>, so it comes from n-dimensional space. So imagine of A as this vector (A1, A2, blah blah blah blah up to An), and each of these values, it basically describes the audio signal, so it represents an amount, so it contains an amount that represents the audio signal of your video or your show. And then the B, in this case, for instance, in this example you can see that the B is then equal to, for instance, 1.2, or B = -1/2. So you can see that B = 1/2, which basically is offensive saying that B = 0.5, or B can be equal to -1/2, which is -0.5. Now then it says then B * A, which basically means multiplying our scalar beta by the vector containing the audio signal A, so this B * A is perceived as the same audio signal but at the lower volume. Now why lower? Because you can see that B = 0.5 or -0.5; it means that once you take all these elements of your A and you multiply it with the number that is smaller than one, in this case 0.5, then all these numbers will decrease, which means that also your audio volume will decrease. So let me actually show you an example. So let's say our talk show is very short, and you know the audio variation is very low. You have a vector A that is quite small; it comes from a three-dimensional space, so R<sup>3</sup>, and it has numbers like 3, 6, and then 5, so a 3x1 vector. And then we have our audio adjustment scalar beta, which is equal to 0.5. Now when we take the beta, we multiply it by our audio signal, then what we do times this clear, so times what we are doing is…

That we are simply taking all the elements of our a, so three, six, and five. And what we are doing is that we are multiplying it by 0.5, 0.5, and 0.5, or you can also say 1/2. So what this is equal to is that 3 * 0.5 is 1.5, 6 * 0.5 is 3, and then 5 * 0.5 is 2.5. And you can see that all these numbers, 1.5, 3, and 2.5, they are smaller and specifically two times less than all the original values in the um original audio. So original audio is a, which was 3, 6, and 5, and the new audio, the the scaled one is, so audio scaled, so B * a is equal to 1.5, 3, and 2.5. So you can clearly see this transform where uh this element three is larger than 1.5, 6 is larger than three, and then the last element five is larger than 2.5, which means that this audio audio is much at a higher volume, so the volume two times higher than this audio. So this is basically the idea of uh applying scalar multiplication to our audio pre-processing. I will leave the other example to you that will show that when your scaler is equal to minus 0.5, you again will end up with a lower volume, only that time the volume will be much much lower than the original one.

So now that we know how we can perform scalar multiplication in theory, as well as we have looked into an example how we can do it in terms of the numbers and multiplying them, and we have also seen uh applying scalar multiplication in practice, uh so we have seen in this audio processing stage the uh multiplication process, we are ready to look into the visualization of it. This will help us to get a better understanding on uh what exactly happens when we are scaling different vectors. Let's look actually in the following example. So let's assume we have a vector. Oh, let me remove that. So let's assume we have a vector, and that vector is, let me get a color, this one for instance, a vector a, and this vector a consists of elements one and two. So where does this vector lie? The vector is with um one, so here in our coordinate system, this is our x-axis, this is our y-axis, and here we got uh let me actually pick another color, let's say black one, and then we got one and then two, right? This is two, this is one, so it is this one. So the line that we get here, it is this one. So this is our vector a.

Now let's assume I want to multiply my vector a, so I want to scale my vector a by a constant three, by scalar three. So I have a scalar, let's say I call k, and this k a different number, let's say k is equal to three. So what I wanted to do is to perform a scalar multiplication, so I want to obtain k multiplied by a, and we learned that this is simply = 2 three times and then one two, and then this is equal to 3 * 1, 3 * 2, which is equal to 3 and then six. So let's also visualize the scaled uh vector. So let me pick this yellow color, this will be our scaled vector. So we have that scale multiplication and we are going to visualize that. So we have three and six, so this is three, one, two, three, and this is six, so we have this point. So you should already see what is going on here. So we got 3a here. So you can see that this part is our vector a, and this longer one is 3a, and even visually you can see that this longer vector is simply the three times of the shorter vector. So we got this, and then if you add on the top of this the same three times, you will then end up with the original, so scaled version of that. So basically this is a, this is a, this is a. We combine three different, so we scale a three times and we simply get a three times longer version with the same direction. So you can see that when we are scaling, even visually it makes sense. So we are scaling our vector a three times and we are just getting that vector. So we are transforming, let me remove this. So basically we are taking this vector and we are scaling it up to this point. If I would do it only two times, then it would be something like this, or one and a half times it would be something like this, or only half of it. So now this should make much more sense. Let us actually do yet another example to uh make sure that we are clear on these visualizations, because we are going to make use of it when uh looking into this idea of linear combination in a span.

Let's say we have a vector b, and this vector b has elements zero and three. So let's visualize and uh plot this vector. So it contains elements zero and three, so zero and three. So this is the x element and the y element on the y-axis. We can see this is three, which means that our vector b is this vector. All right, perfect. So this is our b. Let's now multiply, so scale our vector b by scalar two. So let's say we want to get 2 times b. So what is this amount? This is equal to 2 * 2 * and I'm simply taking each of those elements, zero and then three. So this is then equal to 2 * 0 is 0, and then 2 * 3 is equal to 6. So this is my new scaled vector, 2 * b vector, this one. So let's visualize this. The x-axis value is zero, so we are still here, and then the y-axis value is six, so what is six? This thing. All right, so you already should see that this is very similar what we had before. So this is 2b. All right, so this all uh should make sense. Uh also we learned as part of the um high school when visualizing different plots. So this is quite similar to this idea of having y is equal to x and then scaling it, getting like y is equal to 2x. So in this case only we know exactly where the vector starts and ends uh so we have a much more specific definition instead of having all this infinite number of points on the line, but the idea stays the same. So we are taking this vector and we are then scaling it two times, so we get 2b vector. And I could do the same, only instead what I could also do is I could do like uh 0.5 or 1/2 * b, so I take the half of it, which means I would get this vector, or I could multiply it with minus one, so minus -1 * b, so I would scale with minus one, and then I will simply get the negative version of my original vector, so this thing. This would be -b or -1 * b. So this is basically the idea of uh scalar multiplication when visualizing it in our coordinate system, Cartesian coordinate system. And now when we know all this, we are ready to move on on this idea of linear combination. So let's now talk about another super important topic which is the dot product and its applications.

So we are going to talk about the dot product, and we are going to uh make this relation between the dot product and the distance, something that we have learned as part of the high school. So the dot product, we are going to understand it, we are going to plot the vectors, we are going to perform this dot product in examples, as well as we are going to to look into the properties of the dot product and the inner product concept. We are going to see the relationship between dot product and inner product. We're going to talk about the Cauchy-Schwarz inequality and the geometric interpretation of it. So those are really important concepts, um the Cauchy-Schwarz inequality, um the intuition of it as well as the implication of it, um specifically in machine learning. And then we are going to talk about the vector triangle inequality and uh in what cases also the norm is equal to zero. And finally we're going to define the angle between vectors. So when it comes to the dot product, uh the dot product is also known as the scalar product. It's a fundamental operation in linear algebra. It combines the two vectors to produce a scalar. It's a way to measure how much one vector extends in the direction of the other, so providing insights into the geometric and the algebraic properties of the vectors. So when we are using this operation of the dot product, you can see it as an operation that we apply to vectors, and in this way, once we look into the geometric uh side of it, so we will visualize the vectors and we will see the geometric insights behind it, it will all make much more sense, because this will relate back to our geometric interpretation, this idea of Pythagorean theorem, because they are highly related, this idea of distance and how by using the dot product we can relate all these different concepts.

So let's first formally um define the idea of uh dot product. So the uh dot product of a vector v with itself gives the square of the length of v. So earlier when we were looking into the prerequisites of the course, I did briefly mention that we need to know this idea of length of a vector or length of a line. We said that we denoted this by this concept or by this notation of two straight lines and then the name of the vector and then another two straight lines, and this is what we refer as the length of vector, so the distance between two points or the starting point and the end point, point where our vector is described. So if we have for instance this vector, then if we take this point and we take this point, then we can use that in order to calculate the distance of this vector, and the distance of this vector is defined by this notation. So that's exactly what we have here too. So you can see that we are saying that the dot product is actually this v, so we take the vector we multiply with it itself, we are seeing this is the dot product of a vector and it is equal to the square of the distance or the length of the vector. So the uh in the left hand side we have this notation of the dot product, we just use this dot as usual with any multiplication or any product, and we are saying this is equal to, and we take the length of a vector v and we square it, and this gives us our dot product. Very soon we will see when uh applying it on a uh more general terms what this idea of dot product is when we apply it to vectors just in general.

In a two-dimensional space, when we have each vector represented by two numbers, let's say vector v which is equal to x and y, so we have on the first element the uh x coordinate of our vector in our Cartesian coordinate system, and the second element, so this one is the y, so from y-axis on our Cartesian coordinate system, then the dot product and the length are highly related as usual. So we have here v the vector and then again, so dot product of vector v is equal to x² + y². So this is the length of this vector and it is equal to then the square of this length. So uh let's actually plot this to see it in our coordinate system such that we can relate this back to the Pythagorean theorem and we can see how it is possible that the dot product, the v * v which is equal to the square of the distance is actually equal to x² + y². So let's assume we have this vector v and this vector v is represented by this x and the y. So we are in the R², so this is our x-axis, this is our y-axis, and in here our vector v can be represented, and I'm just taking random some random point and I'm saying this is my x and this is my y, then this is simply the vector v. So this x and y, this can be random numbers, I just want to keep it general and to make it very similar to the definition what we just saw. So we are going to relate this back to the Pythagorean theorem and we are going to prove how Pythagorean theorem is related to this vector and that the dot product of this vector, so v * v is actually equal to the square of the distance of this vector v and then on its own this is then equal to x² + y².

All right, so the first thing that we can do is to check and use a Pythagorean theorem in order to prove that the length of this vector and uh so the distance from this point to this point is actually equal to x² + y², and to be more specific, square root of x² + y². So when we look in here, one thing that we can see is that this side which is the side that we can have here in order to form our triangle, so I want to take the beginning uh where my vector starts and I want to see what is that that we have in here. So this side, given that this is zero and this is x, this means that this entire distance is just x, and of course the other way around holds as well. When we look at the horizontal side, this length, so this is the y coordinate of this point which is y, so we see here that it is equal to this part. So we see that we are dealing with um triangle where so this is our vector v, and we can see that this this side is x and then this side is y. Now we also know that this angle is 90° using the Pythagorean theorem which states that in this type of triangles when it is a right triangle and we have x and y as the sides, we can compute the opposite side, so the distance or the length of this side which is, let's call it, so in this case we have already the vector, so we can say that the length of this vector, this is how we were defining the length of the vector, this is then the square of it according to the Pythagorean theorem is equal to x² + y², because the Pythagorean theorem says if this is a 90° angle here, we have a, here we have b, then this side which is defined by c, so c² is equal to a² + b², and here the c is basically our vector, and we are seeing the length of this vector. So the length of this vector is basically equal to c from here to make it easier to uh to compare the two. So this means that using the Pythagorean theorem we can already find that the length of this vector is equal to the square root of x² + y², knowing that the vector has those coordinates x and y.

Okay, so now when we are clear on uh what the norm is, let's actually look into what we have here. We are seeing that the dot product v v is equal to the square of the length of it. Now we just show that the length of the v vector is equal to the square root of x² + y². This means that this amount, if we remove this, what we were wanted, what we wanted to prove, so this is equal to square root x² + y², and the square of it. So we just want to prove by making use of all the geometric interpretation and the phenomenon that we already know, including the Pythagorean theorem. So this is then equal to square root with the square here it cancels out, so this two, and we simply end up with x² + y². So we already proved that the dot product that we defined like this, and we are saying according to the definition of the dot product, v dot product of the v, so the vector v is equal to the square of the distance or the length of the vector, that this in this specific case, if we are in the two-dimensional space and our vector is defined by this x and the y, then this dot product is simply equal to x² + y². That's something that we just proved. All right, so let's now apply to an actual example. So here we are seeing that the length of a vector is deeply related to dot products, that's something that we just saw also when plotting it in our Cartesian coordinate system. So let's now look into this specific example. So v is equal to 3, 4. So this is the x-axis, this is the y-axis, and here the proof is that the dot product v v is equal to 3² + 4² is equal to 9 + 16 is = to 25, which means that the length is equal to square root of this dot product and it's equal to five. Now let's actually prove this, this just to ensure that we are at the same page. So we can actually go ahead and plot that in here. So we have that the vector v is equal to three and four. So a few things just to get straight before plotting this. So we know that the dot product can be defined like this, and we said that this is equal to the square of the length of this vector, and we also proved that this is also equal to x² + y² if we are in the two-dimensional Cartesian coordinate system. So here this is our x and this our y, so our x is equal to 3 and our y is equal to 4. Let's actually solve this problem and find it also geometrically to prove what is the dot product and what is the length of this vector. So the vector v is 3, 4, which means that it is 1, 2, 3, so this is three and this is four, so we are here and this is our vector v. So this also means that one of the sides of our triangle with 90° angle, so our x, so this our x and then x is equal to three, so x is equal to three and our y is equal to four, so it is this one, this is our y. So now we want to find what is this length, so we want to find what is the length of this vector v. So the length of vector v is then equal square root of, so it is equal to x² + y² according to Pythagorean theorem, and then this is equal to just filling in the numbers, 4² is 16, and then this is 9 + 16 is simply equal to 25, so this equal to 5. So in this way we can see that a we are proving that the the length of this vector v is equal to 5 if the sides are three and four, which means that the vector has this coordinates 3, 4, and at the same time. So in this way we already see a few things. So we see that the length of the vector v is equal to 5, that's something that we already proved, and we also saw that the dot product is equal to the square of the length, which means that it's equal to 5², which means that it's equal to 25. So and that's exactly what we see here, it says that the dot product of v is equal to 25 and the length of that vector is equal to 5.

The dot product of the two n vectors A and B is defined as A<sup>T</sup>B, which is equal to A<sub>1</sub>B<sub>1</sub> + A<sub>2</sub>B<sub>2</sub> up to A<sub>n</sub>B<sub>n</sub>. So this is basically the generalization of the definition that we just saw before. So we go from the um from the description of working with the vector v and the distance to the more general one when we are now dealing with n-dimensional vectors. So now we are no longer in a basic two-dimensional space, but we are in an n-dimensional space, and we have these two different vectors A and B. Think of A as this vector A<sub>1</sub>, A<sub>2</sub> up to A<sub>n</sub>, and then think of this vector b as B is equal to B<sub>1</sub>, B<sub>2</sub> up to B<sub>n</sub>, and the dot product of this two we are writing as A and then this small letter T which means transpose, and we will see something very soon, and we will also be discussing the topic of transpose later in the next module, two multiplied by B. So this A<sup>T</sup>B or A transpose B is um a bit more advanced way of writing down this idea of dot product, and the dot product of these two n vectors A and B is defined as this formula. So we are basically cross multiplying the two, we are taking this element from A and we are multiplying it with this one, then we are taking this element of A and

Multiplying with this element of B, so we are basically taking all the corresponding elements from these two different vectors, and we are multiplying them together and creating this uh combination and adding them up to create this single value to represent these two vectors with the single value, which we call Dot product. So this notation A<sup>T</sup>B, um which is a common way of noting and denoting this dot product, this is just one way of describing this dot product. There are also people who will write down this, so you see that this uh triangle type of braces and then the first vector and then the second one and then this one. So you will notice one thing that will be uh quite different in here versus in this definition, because in here we were discussing the uh dot product of a vector, so we were talking about exactly the same vector. So here we have V and V, and we were just saying let's take the dot product of a vector with itself, and this was related to this idea of length. When we go from one vector to this idea of multiple vectors, so now we are creating no longer the dot product of a single vector, but we are defining the dot product of two vectors A and B. The definition changes, as you will notice, but that's important to differentiate the two and to understand that one is actually not the other one, and to um to be able to see the difference between the two and not to confuse them. So uh in the upcoming presentations we will use this notation of A<sup>T</sup>B much more often than this uh this second notation, but if you see this, this uh you can just know that we are talking about the same thing, which is that the dot product.

So when the N is equal to one, then the then this inner product or the dot product simplifies to the multiplication of just two numbers, because we no longer have vectors, but we just got a single value for a and a single value of B. So let's assume that the N is equal to 1, and then a is equal to 2, and then B is equal to 3, and then the dot product is simply the A<sup>T</sup>B uh which is equal to 2 * 3, which is just 6. So uh when it comes to this uh idea of transpose, I will explain this in a bit, and we are going to uh describe this in detail as part of the next unit, so I won't go too much into it, but it's always a great idea to to know um also for this module um at a high level this idea of transpose, this T is basically referring to this, so transpose.

To understand this idea of dot product, the dot product is an operation that takes these two vectors, vectors A and vectors B, and returns a single number, a scalar. And the single number is uh computed by calculating the sum of all these individual values. So it takes all these elements from one vector, then the other one, so A<sub>1</sub> and B<sub>1</sub>, A<sub>2</sub> B<sub>2</sub>, A<sub>n</sub> B<sub>n</sub>, takes the value from one vector with the corresponding one from the other vector, multiplies them to each other, and then add them all up, and then this single value is created, which will refer as the dot product. Geometrically, it represents this product of the vectors' magnitudes and this cosine of the angle between them. So um the we already know from high school and also from trigonometry and geometry this idea of angles and what is the cosine of an angle, so I won't go too much into into the detail of it, I will assume that you know what a cosine is. So the dot product is used to determine the angle between the two vectors and also to check whether they are orthogonal or not, because uh we have a property that relates to dot product directly to the orthogonality of the two vectors. And knowing how to calculate the dot product of the two vectors will help us to check whether the two vectors are orthogonal or not. And just in general, understanding this concept of dot product is super important because we use dot product in machine learning; we use dot product in uh generally in AI; you will see that dot product appearing also when you when we are talking about attention mechanisms or calculating this uh attention scores as part of multi-head attention in Transformers. We also use this idea of dot product, scaled dot product, so it's really fundamental and essential to understand this concept of dot product for matrix multiplication, but also in general to apply linear algebra in different science-related domains.

Let's now look into a specific example to um understand this concept. The dot product of these two uh vectors, in this case we have three-dimensional vectors A and B, the dot product can be calculated as follows. So you can see that we are taking the one, this is the A<sub>1</sub> and we are multiplying it with B<sub>1</sub>, so it is the this one, so B<sub>1</sub>, B<sub>1</sub> is equal to 4, A<sub>1</sub> is equal to 1, so it is this one; we multiply the two, and then we go to the next sum. So here I will write it down actually underneath to make it easier. So this is A<sub>1</sub>, this is A<sub>2</sub>, this is A<sub>3</sub>, as you can see, and this is B<sub>1</sub>, B<sub>2</sub>, and B<sub>3</sub>. So taking A<sub>1</sub>, so taking A<sub>1</sub>B<sub>1</sub>, adding A<sub>2</sub>B<sub>2</sub>, two, and adding A<sub>3</sub>B<sub>3</sub>, we are getting the dot product, which is equal to 4 - 6 + 5, which is equal to 3. So if you do the calculation or you use a calculator, you will find out that the dot product that can be described by this formula, which is very similar to what we have seen in here, only n is equal to 3, we are getting that the dot product of the vectors A and B in R<sup>3</sup> is equal to 3.

Let's now look into another example where we will do all the calculations manually step by step. So we have these two different vectors, we have Vector a and Vector b. Let me also here the arrows, and I want to calculate the um dot product a · b. So the A<sup>T</sup>B, the dot product A<sup>T</sup>B is equal to, and we know that if we are in the N three-dimensional space, so we are in the R<sup>3</sup>, we expect that we will need to just use the cross product three times and we need to add them up. So we know that it's equal to A<sub>1</sub>, let me use the black color, so A<sub>1</sub>B<sub>1</sub> + A<sub>2</sub>B<sub>2</sub> + A<sub>3</sub>B<sub>3</sub>. Now what is A<sub>1</sub>, what is A<sub>2</sub>, and what is A<sub>3</sub>, and what is B<sub>1</sub>, B<sub>2</sub>, B<sub>3</sub>? So A<sub>1</sub> is equal to -1, B<sub>1</sub> is equal to 0, A<sub>2</sub> is = to 2, B<sub>2</sub> is equal to 1, and A<sub>3</sub> is equal to 2, this is 3, and then B<sub>3</sub> is equal to -3. Now what is this amount? -1 * is 0, 2 * 1 is 2, and then 2 * -3 is -6, and what is 0 + 2 - 6 is -4.

Let's now understand what is this application of dot product and how it can be used. So the dot product between these two vectors gives us an idea, uh a way to measure their similarity. So it combines the magnitude, the length of the vectors, and the cosine of the angle between between them. In simple terms, it tells us how much one vector extends in the direction of the other one. Now we know how we can calculate it once we have the vectors; we we understand the notation; we also know this relationship between the um vector and the uh distance; what we mean when we have um norm, when we have a distance; we also know the the length of a vector. I want to combine all this and uh combine this with this idea of cosine, which is something that we learn in our high school as part of our trigonometry and geometry. I want to relate this all to make a sense how geometrically we can interpret the dot product, and it will all uh make much more sense. So the dot product between the two vectors gives us a measure of their similarity. This is something that we are also using a lot in the field of artificial intelligence when we want to compare compare two different uh, for instance, users. So we have one user, and we want to compare the uh this one user to the other user, then we are using this cosine rule; we are using what we call cosine similarity as a way to measure their similarity. So how similar those two customers are, how similar those two users are, how similar those two companies are, and we are doing that by using this simple cosine idea, which also is related to this idea of norms, this idea of distance, so or the length of a vector and the relationship between two vectors.

Let's actually wrap this up together and um clearly uh indicate what is this relationship between the dot product, the length of vector, cosine rule, and this cosine similarity. So in simpler terms, this uh dot product um it tells us how much one vector it extends in the direction of the other one. So uh when vectors point in the same direction, then we know that the dot product is positive and largest; it means that we are dealing with two vectors that are very similar. But when the vectors are perpendicular, then the dot product is zero, and we say that those two vectors do not share any direction; they are not similar. So when the vectors point in the opposite direction, then the dot product is negative; we are saying those two users are really um negatively correlated, so one is the opposite of the other one, so to say.

Let's actually plot this and see it um and bring them all together. So uh let's say we have a vector s and we have a vector r, and we want to see how similar the two are. So we have here our vector s and here we have our vector r, and those two vectors come together and they um create this angle, and we are referring this angle by angle θ. This is just a way to refer to our angle; those are all things that we learned as part of high school through geometry, so they should all seem very familiar. So uh from the geometry we know that if we have this triangle and here we got this um angle which is θ, and here we got uh let's say side c and then here a and then the b, we know that the c<sup>2</sup> is equal to a<sup>2</sup> + b<sup>2</sup> - 2 * a * b * sine of the opposite angle of the c, which is θ. Now let's actually go ahead and apply this to our case when we instead of having just a triangle we got a vectors, so we got a set of three vectors that form this triangle, and we already have seen in many examples previously one way of doing that, and that's actually go ahead and transform this vector uh this uh into vectors, and we can do that if we plot in here this side, then this, so we basically recreate the same triangle, only this time the sides are just vectors. So this is let's say um our vector r, this is our vector s, and then this is obviously the r - s, and the opposite side is α. So the vector r and vector s they form the angle α. So when we apply to the cosine rule to this, of course we instead of using vector notation we need to use their distances because the cosine rule is based on the sides, so the length of the sides. So the length of the uh of this side, this one, we denote by our common notation of the norm, which is simply this thing, and then the same holds for this side, which is simply the norm or vector s, and then for the third side, for this one, we then have norm of r.

So if we apply this cosine law to our specific example, we can see that in our case our c<sup>2</sup> is simply our the length of our vector r - s only squared, which means that, so let let me actually write it down for simplicity that part, so a is then in this case my vector r, and then the distance of it obviously, so the length of the uh vector r, and then b is the length of vector s, and then the c is simply equal to the length of r - s. Then here I will just change my angle and I will call it θ, and let's now apply the cosine law to our actual example, which means that we got r - s and then squared is equal to, and then we got r squared and then plus s squared, so the length of the vector s and s<sup>2</sup> - 2 * r * s * cosine of θ. So one thing that we can recognize here very quickly is that here we are dealing with the s<sup>2</sup> of a length of a vector r - s and we have learned from the previous definition of a dot product of a vector with itself, so the dot product of vector v is equal to the square of a distance, so this means that here we already got the square of a distance of the vector v, so this is something that we have because we have r - s and then squared, this is exactly the same as the vector v vector v and a squared with its length, so the length of vector v squared is the same as the length of the vector r - s<sup>2</sup>, because in this case the v is equal to, so the vector v is equal to vector r - s. And uh from this we can see that this means that here we are dealing with the dot product of r - s. So let me actually write this down in here in our example, so this is basically the right hand side of the formula that we saw before where it was saying that vector v dot product of vector v, so this is equal to and then vector v the distance of or the length of it and then squared. So in this case the vector v is = to r - s. This means knowing this we can actually make use all that formula and say that here we are dealing with the dot product of r - s, which means that we can write that this can be rewritten as r - s and then times r - s. So in the left hand side we will be starting working with the dot product with the vectors, while in the right hand side we will still keep our lengths. So here we have the length of the r vector, so r<sup>2</sup> + s<sup>2</sup> - 2 * r and then s and then times cosine of θ. Now let's actually open those parentheses, so this gives us r and then r basically the same dot product of vector r and then - sr, so dot product of vector s and r and then - s actually + s and then s and then - s * r, so I'm basically opening the parenthesis of the uh above part this left hand side to simplify, and then in the right hand side I will just copy paste this, so r<sup>2</sup> + s<sup>2</sup> and then - 2 * r forgot here arrows and then s and then cosine of θ. Let me actually go ahead and remove all these parts to open up some space for us, and there we go right. So let's actually go ahead and simplify what we got here now. What do we see here? We see here that here we got one dot product, and we know that the dot product of a vector r r is simply the distance of that vector and then squared - here we got, so we will just keep it like that, we are taking the dot product of vector s with vector r, we are adding here, here we again got a dot product of a vector with itself, which is equal to the the the length of that vector s only the square of it using exactly the same logic that we just saw, the dot product of a vector, and then here minus and I will keep the dot product of these two vectors s and r together, so this is then equal to and I we just take over this part exactly as it was, so r<sup>2</sup> + s<sup>2</sup> and then - 2 * r and then s and then cosine of θ. Now uh we already see couple of things that we can cancel out, which is great, so we see here that we are dealing with um the length of a vector r and a squaring, something that we see here too, so we can safely cancel those out. Another thing that we can see that we can cancel is this, so again the length of the vector s and squared, so this also go away, and we end up with this much more simpler formula, so we got -s and then r and then -s and r and then this is equal to -2 and then r and then s then cosine of θ. Now we already see that this we can combine, we can say this equal to -2 times and then the order doesn't matter of course in the dot product, so r and then s and then this is equal to -2 * and then the distance or the length of the vector r and then the length of a vector s and then times cosine θ. Okay, perfect. Now we can get rid of -2 to, so let me clear some space up to see what we end up with, so we are left with the following formula, so we see that r and then s, so the dot product of these two vectors is equal to r, so the length of the vector r multiplied with the length of vector s multiplied by cosine and then θ. So this is basically what we end up with and exactly what we wanted to prove; we wanted to see that the dot product, so this is the dot product of vectors r and s is equal to the length of the vector r multiplied with the length of a vector s multiplied by the cosine of this vector that they are forming, so cosine of θ.

Now why is this important? This means that we can also take this cosine, bring it to the other side, and we can use that to describe the cosine with this uh dot product as well as the length of these two um uh vectors. So from here we can say that the cosine of θ is equal to the dot product of the vector r and s divided to the length of vector r multiplied with the length of vector s. So using this we can then calculate the cosine similarity between two different vectors, and this then helps us to mathematically compute what is this measure of similarity between two different entities, whether do we want to compute similarity between two customers, whether do we want to compute the similarity between two users, similarity between two uh profiles of uh clients, so so this can be applied really to do very large range of applications, but the idea is that in this way we have proven that we can use this algebraic concept, the linear algebra concept of the dot product in order to compute a similarity measure that will help you to compute the similarity between two different entities. And this is a very common application in machine learning, in artificial intelligence; the cosine similarity is one of the most popular uh similarity measures that is used in data science, in machine learning, in artificial intelligence, whether it's in the clustering algorithms like K-means, whether it's in different sorts of machine learning algorithms or just in general to compute the similarity between uh two uh reviews that customer leave, similarities between two profiles of customer, similarity of two different um uh users, for instance two users that are watching movies and you are building a movie recommender system, you can use cosine similarity to compare the two users' behavior in order to improve the quality of the movie recommender, and the applications are so versatile that um my takeaway uh would be to just uh keep this cosine similarity in mind, the formula of it and how it is related to the dot product, because uh this is something that will be very useful uh for for your uh journey in applied sciences.

Let's now talk about the idea of inner product and how it relates to the idea of dot product. So by definition, the inner product generalizes the dot product to more abstract vector spaces, which can include spaces with complex numbers. So it retains these key properties like linearity, you know, linear independence and notion of angle and the length that we saw before, the norm and how we related back to cosine similarity for much more complex vectors. The inner product involves uh basically conjugating one of the vectors, and in real vector spaces the inner product is simply the dot product. So if we have vectors that are from uh real spaces, so they are from R<sup>2</sup>, R<sup>3</sup>, then in this case the inner product is

Exactly the same as the dot product. So let's now move on towards uh one of our final uh topics as part of this unit, and this is a very important topic, um, very often referred to as Cauchy-Schwarz inequality. This inequality and the law states that for all vectors X and Y in an inner product space or dot product space, the absolute value of their inner product is less than or equal to the product of their norms. Mathematically, we can express this as the following expression:

So where this part it simply represents this dot product of X and Y, and here, as we already saw before, the ||X|| with this specific notation it refers to the length or the norm of our vector X, and then this one, the norm of vector Y. So this definition should, and this um, and this way of referencing the norm and the length as well as the dot product should seem very familiar to you.

So what we're saying here is basically if we have a vector X and the Y and we have a dot product of X and Y, if we take the uh magnitude, then this will be smaller than or equal to the length of the vector X and the length of vector Y. The intuition behind this law is that it tells us that the absolute value of this dot product between two vectors cannot exceed the product of their lengths. So basically, if we have a vector A and vector B and we are computing the dot product between them, so A.B, then this amount will always be smaller than or equal to the length of vector A multiplied by the length of vector B. This is all what this law is about. Oh, this inequality, it tells us that the absolute value of the dot product between two vectors cannot exceed the product of their lengths.

So let's actually look into and understand why does this even make sense. This idea, this inequality, Cauchy-Schwarz inequality, it is about the limits of similarity. So what is this level, the maximum level that we can have when it comes to the similarity of two entities? For instance, if we got two um users and we want to understand how similar those users are, this inequality helps us to understand what is this maximum level that we can have when it comes to the similarity of these two people. What is this maximum similarity measure that we can have based on the characteristics of these users? So no matter how similar the two vectors are, there is an upper bound to this measure based on what kind of vectors we are talking about, and this uh maximum amount, this maximum measure is defined by their lengths. So it confirms that this uh cosine of an angle that we just saw, because we saw that the cos θ is equal to is exactly what we got in here. So we saw that cosine of θ is equal to the dot product of vector R and S divided by the length of them, and that's exactly what this Cauchy-Schwarz is basically saying. It's saying that it confirms that this cosine of the angle between them, between these two vectors, part of this dot product calculation, it cannot exceed one, which feeds to our understanding of the cosine because we know that cosine is a volume that should be between minus one and one, to be more specific, minus one and one included. And knowing this, this helps us to get a limit on this idea.

So how can we prove the Cauchy-Schwarz inequality? This comes down to this formula that we have already proven. So let us actually go ahead and prove this Cauchy-Schwarz inequality. Quite straightforward. So we have just learned and we have actually proven that cos θ is equal to R.S if those are our two vectors, and then here we got the length of the vector R and then the length of vector S. Now we know that a cosine is a measure that will always be between -1 and 1. So this means that we can use in order to prove this. So knowing that this volume will always be smaller than or equal to one, it means that we can just take this and we can bring it here. We know that the lengths are both positive amounts, so it means that I can just multiply both sides by this amount, that means that I will be left with R.S, so the dot product between the two vectors that will be smaller than or equal to, and then R, the length of this vector multiplied by the length of the vector S. Now, as you can see, this is exactly the Cauchy-Schwarz inequality that it says because it says that the absolute value of the inner product is less than or equal to the product of their norms.

The Cauchy-Schwarz inequality is a fundamental principle in mathematics, but also just in general, it's used in so many applied sciences and it's applied across various domains including linear algebra, physics, but also machine learning and in finance. So this helps us to establish a limit on the correlation or this idea of similarity that can exist between two vectors in any inner product space. So if we got two different companies and we want to understand how similar they are performing, we, by knowing their performances and their measures, we can uh then have a limit on how similar the two companies can be or how similar the two users can be based on their previous history, and this all is being done by using this idea of Cauchy-Schwarz inequality.

Now why does this matter? So understanding this inequality gives us deeper insights into the vector spaces, especially when working with this high-dimension data, which is quite uh often the case when we are dealing with Transformers, with large language models, with uh deep neural networks, or um we are working in finance. So this is a super uh important concept when it comes to applied sciences. So it ensures that this um validity of operations when it comes to the mathematical side and the transformations we perform on our data, on the practical side, that those two are coherent and we ensure that we are safe in making different assumptions and we provide in this way the safety net for um our design of the algorithm. We ensure that our algorithm is performing well. So every time when we are uh creating a program or an algorithm for a specific case, then using this inequality we can basically add use cases, we can create test cases, and we can add these constraints to our algorithm, and we can design our algorithm in a more intelligent way, and we can ensure that we won't have cases when the uh limit goes above uh this threshold that we already got based on this inequality. And this is a very important application in a practical sense when it comes to applying linear algebra to applied sciences.

Now, uh the final question that I will leave you with is the case when the norm is equal to zero. So we briefly touched upon this point and we spoke about this idea of perpendicularity, and in a one specific case when the norm of a vector v equals zero. So the theorem says that the norm of a vector v equals 0 if and only if v is the zero vector. So mathematically, this means that the distance of the vector v is equal to zero. So basically, we are dealing with a zero vector, so v is equal to 0. This means that the only vector with no magnitude pointing in no direction is a zero vector. Only in that case we will get a norm equal to 0. This concept is super important to understand this uniqueness of a vector and this different foundational principles and properties that will also see appearing in the upcoming lectures and in the upcoming lessons.

This is all for this module and for this unit. Uh, we have learned a ton of uh important stuff when it comes to the vectors. So we have looked into the uh from very scratch the idea of vectors, the norms, the distances, the inner product, the dot product. We have looked into the geometric interpretation of it, we have visualized it, we have looked into different properties of vectors, the spaces. We have also looked into the properties of the dot product, we have proven them, we have looked into tons of examples. We have understood how we can plot vectors, how we can understand and relate the idea of vectors to this dot product and the dot product to the cosine of two vectors and the cosine of an angle and the relationship between two vectors and the similarity between these vectors, how we can use and apply that in machine learning, in artificial intelligence. We have kind of made those relationships there, and um we have also looked into fundamental concepts including the um uh scalar vector multiplications, operations with vectors, adding them, subtracting them, and um we have also looked into this idea of linear combination, span of vectors as well as linear independence. We have defined linear independence and we have also spoken about uh and provided examples for vectors that can be uh marked and defined as linearly independent and also examples when our vectors were linearly dependent. Thanks for uh staying with me uh so far, and I will see you in the next unit.

Welcome to module one of this new unit when we are going to talk about about matrices as well as linear systems. So those are all fundamental concepts that you will see time and time again when applying linear algebra, not only in mathematics but also in applied sciences like data science, artificial intelligence, when training different machine learning models and trying to see what is this mathematics behind machine learning models, different optimization techniques when you want to solve different problems using linear algebra. So in this first module as part of foundations of linear systems and matrices, we're going to introduce this concept of linear systems, and then we are going to talk about the general linear systems. We are going to uh see this common labeling, all the coefficients, this idea of indices that refer to the rows and the columns. We're going to see what is this differentiation between homogeneous and non-homogeneous systems. So without further ado, let's get started. So the linear systems form the uh bedrock of linear algebra modeling this array of problems. Thanks to this advancements in these linear systems and solving it in computing, we can now solve a large amount of problems in a very efficient and a fast way.

So uh the general linear systems can be represented by this uh set of M equations with n unknowns. In the previous unit when we were looking into this uh linear combination of vectors, we saw this notation which was A1 and then we had c1 multiplied, or rather let me keep me uh let me keep the same notation, so we had this linear combination of vectors, so we had β1 and then we had A1 + β2 and then A2 and those are all vectors + A3, so β3 * A3 ... and then βm * Am. This is the notation that we saw before, and we said we want to come up, we wanted to come up with the linear combination of these different vectors A1, A2, A3 up to Am, and then we use that in order to get a sense of whether we are dealing with linearly independent variables, vectors or linearly dependent vectors, and then we also commented on this span that this vectors take. Now when it comes to um the uh vectors and just in general linear systems, we can represent what we had before now in terms of with uh bigger system, so in terms of M equations and with n unknowns. So here what you can see here is that we have M different equations, so we have B1, B2 up to BM, so you can see it in here, and then each of this equation it contains n unknowns, so you can see that the unknowns it stays the same, so the unknowns are those X1, X2 up to Xn. So X1, X2 up to Xn are the set of all n unknowns, and then M equations that you can see in here are all these equations: a11X1 + a12X2 ... and then a1n and then Xn is equal to B1. And here one thing that is really important to keep in mind is that the indexing is what we need to focus on, so we need to keep this one in mind, this aij and this Xi. So this is something that we also spoke about when uh discussing the linear combination of vectors, we slightly uh touched upon on this topic. So let's now dive into this, this indexing and how do we index aij, what are these A's, what are these J's? And here you can see that we have a11 and then a12 and then up to the a1n, and this is in our equation one, and then we have in our equation two, a21, a23 up to a2n, and this a that you see here, those are just real numbers. So a11 can be 1, a12 can be 3, a1n can be 100, and then the same also holds for this B1, for this B2 and for this Bn, and all these values A's and B's they are just real numbers. The only unknowns that we got here are those, so the X1, X2 up to Xn. All right. So what about the indexing now? So we got aij, and as you can see in this case, the first thing that we can see here it stays everywhere the same which is the one, so we got here one, we got here one and up to the point we got here one, whereas the second index, this one, it does change, it grows gradually with one and it becomes, it goes from one to two and up to n. So you can see here that the first index, first index or index I, it goes from one, it doesn't change, it's just one, so it is 1, 1 and 1. So here in all cases for this equation I is equal to 1, but another thing that you can notice here is is that the index 2, unlike index I, so the second index which is the J, so you see here that the second index is referred to as J, this is a general way of defining the indexes. So here J is equal to 1, 2 ... and then n. So basically the I doesn't change in the same row, but the J stays the same, and then it is I 1, 2 up to n, but the set is the same, so it is it contains all these different elements here, so 1, 1, 2 and then n, but it contains all these different real numbers going from one till n because we are combining and we are creating this combination, the sum of all these values a11 and then X1, a12X2, a1nXn, and another thing that you can also notice here is that here with the second index, so with this J, J is equal to one, then here the X's corresponding index is also one. When the J is equal to 2, then the X's corresponding index is also two, and then here the same story, and you will notice that while the coefficient contains two indices, 1, 1, 1, 2 or 1, n, which are the two indices for the coefficients for the unknowns, we got just single index which goes from one till n. So basically for A's, for the coefficients, so I, let me write with the right color, so I can be 1, 2 all the way to M, whereas in case of J it can be 1, 2 all the way to n, and the indices are basically used to help us to keep track of in which row we are and what is the um variable that the coefficient belongs to, because knowing this second index, this helps us to understand that we are dealing with a coefficient that corresponds to this first unknown, the first variable X1, and then the same holds in here as you can see in here and in here we are dealing with the same variable X1, therefore the second index, the index J is then the same both in the first equation and in the second one, in both cases it's equal to one.

Okay, so now when we are clear on that, let's also understand this high-level concept because you will see this system of linear systems, this M equations and N unknowns appearing a lot, not only in terms of calculating and finding the solution to this linear system, but this actually has a very common application when it comes to um running regression, linear regression specifically. And one thing that you can notice here is that here we got also this B1, B2 up to BM, and you will notice that here the index also uh goes from one, but then this time to M. So when it comes to the rows we have M rows or M equations, therefore we also expect when it comes to counting from the top that at the bottom we will see an M, whereas if we count from this side, so kind of like imagine it like a column, then we see that it goes from one till n. So those are common observations and reference to um number of observations and number of uh features that you will see in your data when dealing with data analysis or modeling data. So just this uh just keep those things in mind, this uh abbreviation of M and then n, M equation and unknowns, because this will become very handy, and the same also holds for this indexing, just to keep in mind that this I and this J, what those indices are and how, for instance, the first, you know, the I, the first index changes when we go from up to the bottom and how the second index J goes and it changes when we go from left to the right when we go through the columns, but we are going to see this also in the upcoming slides, so uh we can we will have time to practice it. So um this is what we are calling a coefficient labeling, the coefficient uh aij. So this thing in a linear system, they are labeled where the first index represents the row and the second index denotes the column. So when we see aij, we know that this refers to the row and the J refers to the column. So this is something that we use in order to understand where exactly in our matrix, something that we can we will see very soon, where exactly our unit or our uh member that is part of our matrix, where exactly is that located, in which row and in which column. The systematic labeling is super important because this helps us to keep the structure and this helps us to understand uh what does this uh coefficients are present, what what is this row that it belongs and what is the column it belongs, so for which equation and for which unknown we have already solved the problem such that we can know what this uh coefficient represents.

So before moving on onto the actual linear systems and this definition of matrices, let's quickly understand this distinction between homogeneous and non-homogeneous because this will help us to also get an understanding how we can solve a system of linear systems. So a system is homogeneous if all the constant terms Bi are zero; otherwise it's non-homogeneous. So identifying this helps us to really understand the nature of the solution set that we need to get and to understand what kind of strategy we need to to use in order to solve this problem. Now what do I mean by Bi? We just saw that we had this system of M equations with n unknowns, and we saw that that we have in the right-hand side this B1, B2 up to BM, which means that we had this M different equations with n different unknowns, and to find a solution to the system it means finding this values corresponding to X1, X1 here, X2, X2, Xn, so basically finding the set of X1, X2 up to Xn that solves this problem, and for us to know how to solve this problem we need to know whether this B1 is equal to zero or not, this B2 is equal to zero or not, and then this BM is equal to zero or not. This is very similar to this idea of solving any sorts of um problems that contain unknowns. For instance, if we have 3x is equal to let's say 5, solving this is entirely different than if we know that 3x is equal to 0.

So this is a super simplified version of course, but the idea is the same. Knowing that this B1, B2 up to BM, this R zero, this gives us an idea how we can solve this problem. And later on, we will see this distinction between non-homogeneous and homogeneous system. And whenever this BS—so whenever this B1, B2 up to BM, whenever these BS are zero, then we are saying that the system is homogeneous, and we need to solve a homogeneous system; otherwise, we are dealing with non-homog system. So this means that the bis are not all zero.

Let's now move on to the second module, which is about the matrices. So we are going to define the Matrix; we are going to see the definition of it as well as the notation, the IDE of rows, columns, Dimensions, uh, some of which we have already touched upon, but we are going to uh go into the depth of it. We're going to learn properly as well as we are going to see many examples. Then we are going to talk about Matrix types. So here we will talk about identity Matrix, diagonal matrices, and also special type of matrices like matrices containing only zeros and only ones.

So by definition, a matrix is a rectangular array of real numbers that are arranged in rows and in columns. For example, an M by n Matrix a can be represented as follows. So, so let's look into this definition and this reference to Matrix. We call this Matrix or Matrix a, and every Matrix it can be described by this rows and columns where we always have this uh way of describing this Matrix. Always should be defined by the number of rows and number of columns. So this is super important, and let's look into this specific Matrix. So we have a matrix a, and all these values, they are members of this Matrix; they form the Matrix. And we already saw this labeling of a<sub>ij</sub> where we said that I is referred to the row. So you might recall that those all these equations that we got—so this horizontal lines where I was equal to one, I it was equal to two, I was equal to three up to the point of I was equal to M—and then we had this J, so this thing, and then J was referred to the columns, and we had J was here, here one, and then two, and then three up to the point of n. So 1, 2, 3, and N. This is exactly what you can see here.

So in this Matrix, we got all these elements: a<sub>11</sub> is a number, a<sub>12</sub> is a number, up to the a<sub>1n</sub> is a number. Those are all real numbers. And one thing that you can notice here is that here we got a<sub>11</sub>, so this is our first row and First Column. Here we got a<sub>12</sub>, this is our first row and second column, and then we got up to the point of a<sub>1n</sub>. Actually, let me just write this down even at a bigger scale such that I can make more notes. So let's assume we have this Matrix a, and this Matrix a if I'm bigger, and we got all these different elements. So we start with our first row, and here we have a<sub>11</sub>. So here the row that I will write with, let's say, we blue, the row is equal to one, and then the column is one. So this is Row one, this is Row one, row one, and this is column one. Let me write it, wait, right, this is column one, this is column two, this is column three, dot dot dot, and this is column n, and this, this is row two, this is Row three, dot dot dot, and this is row M. So in total, I got M rows and N columns. I will come to this notation that I'm putting here later. For now, let's keep track of the rows and the columns to get a good understanding what this indices were about that we just learned. So every time I will also mention this reference to a<sub>ij</sub> to keep track of this, and also let me write it with the right colors. So a<sub>i</sub>, this is the row, and J, which is the column. So all the elements, I'm just defining by this a because it's just a way to reference a part that comes from a matrix; it's a just common way to write the entire matrix by capital letter A, whereas its members we will write with the um with the lower case a. So this is Matrix, Matrix a. All right. So here in the second row, but First Column, we got a<sub>21</sub>, and then one because it is still in the first column. And then when it comes to this element, we have here a, the row is the first one because we are in the first row, but then we are in the second column, so this one should be two. Then we go on to the next element in our first row, so a<sub>1</sub>, and then three, and then dot dot dot, the last element is an a, as we are still in the first row, it will be one, the I, but then given we are in the last column, the column index of the J will be equal to n because we got in total n columns.

So we are now ready to go into the second row. So here, given that we already have our first element a<sub>21</sub>, this is in our second row and the First Column, so the I is equal to here two, and G is equal to 1. Let's now write down the element in the second draw, second column, as you might have already guessed, I is equal to here one, I equal to here two, sorry, and then uh the J is equal to 2, and then we go on to the next element which is in the second row and the third column, so it's a, the row index is two, so I is equal to 2, and then the column index is three, dot dot dot, and then we got a, as we are in the second row, it is the I is equal to 2, and as we are in the last column, the J is equal to n. Now you might have already guessed when I was writing this down that whenever you are in row and you move on to all the elements in the same Row, the I, so the row index, it stays the same, only you need to uh update the column index. So here, for instance, you got one, one, one, here also one, so all the way down in the same row or one, which logically makes sense because we are in the same row, so the row index should not change, but instead you should change the column index, like here column one, column two, column three, all the way to column n. So those are our columns, dot dot dot. So let me make this distinction, and those are our rows, as you can see. So this kind of mentally helps us to understand why we are writing all the indices over time. Once you practice more with this, this will become more natural very quickly. Remove this. So now I will write the rest very quickly. So as you might have already guessed, we are in the third row, so we have a three, so everywhere I would just write down the aces, so first I'll write down the aces, and then the rows, the row index will stay the same as I'm in the same row, but then I will increase the columns gradually. So we are in the column one and the column two, column three, up to the column n. So now the remaining stuff you can actually write down yourself to just practice. Let's now move on onto the last row and last column. So in the last row, we got a, a, a, up to here, here, and in the last row, the uh row index is M, which means that here I need to have M, M, M, everywhere I need to have M, and then the column index is 1, 2, 3, all the way to n. So this last column is very interesting too. You can see here that we have the opposite of what we have here because in the last column we see that the uh column index is the same, so it is everywhere n. Only the first index, the index of the row, it changes; it goes from 1, 2, 3 up to M, which is of course logical because we said that in the last column, if we are looking it from the perspective of column, so all this values, this A's, so the all the ends, they are logical because they we are in the last column, we are in the same column, but then the row changes. Here we are in the row one, here we are in the row two, Row three, after to row M. Therefore, we have also at the end a<sub>MN</sub>.

Now let's talk about this idea of MN. We said that our Matrix a has M as a number of rows and n as a number of columns, which you can see by the way also here. So we always refer the dimension of a matrix, so the dimension, dimension of Matrix a by these two numbers. So first we always write down the number of rows, in this case M, then as the second element we are writing the number of columns, in this case n. We are always putting this small X in between to kind of emphasize M by n Matrix, and we most of the time use the square braces to Showcase that we are dealing with Dimension, and in this case we are saying the dimension of Matrix a is equal to M by n. So we are dealing with M by n Matrix. This is a common convention used in linear algebra, in mathematics General, but also used in data science, uh, in machine learning, artificial intelligence. So whenever you are dealing with matrices a, it is a common convention to talk about this idea of dimensions, and the idea of Dimensions is super important when it comes to the idea of multiplication, multiplying Vector with Matrix, Matrix with Matrix. So this dot product, Dimensions play a central role in here, so keep this one in mind. Once we uh get to the point of that products, this one will become very handy.

Let's now look into a specific example where we see simple Matrix a. So in this case, you can see that we are dealing with a matrix that has a 2x3 Dimensions. So like we just learned, 2x3 means that we got two rows and three columns; that's something that you can also see here very quickly. So you have a small Matrix on the small matrix; it's really easy to actually count. So you can see that we got Row one and row two, and we got column one, column two, and column three. So this basically confirms these Dimensions. Therefore, we are also saying that we have a 2 by three Matrix, and like usual, we first write down the number of rows and then the number of columns. You can see here that here we have this elements for our Matrix, so a is equal to 1, 2, 3 for the first row, and then our 4, 5, 6 for the second row. So from this, actually, I think it's a good exercise to just uh verify our understanding of IND, and from this um we can write down that for instance all these different elements uh like a<sub>11</sub> is equal to 1, a<sub>12</sub>, which means that we are in the first row and in the second column, so we have this element is equal to two, and then we got a, and then 1, 3, so we are in the third column, so this one is equal to 3, and then a<sub>21</sub> is equal to 4, a<sub>22</sub> is equal to 5, and then a<sub>23</sub> is equal to 6. So this is actually a good way to practice our understanding of indices, our understanding of this Matrix structure, and the understanding of dimension of the Matrix, which in this case is 2x3. So this is yet another different definition of a matrix structure when it comes to the rows, columns, and dimensions. So this is exactly what we just spoke about on our example, and let's just quickly look at the formal definition. So the rows of a matrix are the horizontal lines of the of the entries, while the comms are the vertical lines. So basically it's saying those are—let me remove this—so the rows are the horizontal line, and the columns are those vertical lines; those are the columns. This helps us to form these columns, so col one, colum two, and colum three, whereas this horizontal lines it helps us to create the rows, so Row one and row two, which is a core Matrix operations. So when it comes to matrices, we often perform Matrix additions, Matrix subtraction, but also Matrix um a scalar multiplication of this Matrix, so multiplying Matrix MX with a scalar, and then Matrix um multiplication just in general, so taking two matrices and multiplying them. We are going to look into this concept in detail; we are going to see many examples like before; we are going to dive deeper into this such that we lay the ground on uh to the next module, which is solving a system of M equations with an unknown, so solving this General U linear system. So for the beginning, uh we will be looking into this Matrix operations where we are adding or subtracting matrices.

So by definition, the sum of two matrices A and B of the same dimensions is obtained by adding their corresponding elements. So by taking the element i<sub>j</sub> from both matrices and adding them to each other. So in this case, you can see that Matrix A and B are here, and uh the uh definition says we just simply need to take the corresponding elements, corresponding elements from the row I and the column J, take them, add them, and this will become an element in our final um Matrix. Because when we are adding two matrices of the same size, the result is yet another Matrix. So we will use the Matrix a to add to Matrix B, and this will give us a matrix A + B, and this i<sub>j</sub> simply refers to the indices corresponding to the row and the column. We will look into an example in a bit, and this will make much more sense. And the same holds also for the difference. So by definition, the difference of the two matrices A and B of the same dimensions is obtained by subtracting their corresponding Elements, which means that in order to obtain this Matrix A - B, this is a new Matrix, we simply need to look for each element. So we are going to index them for a row I and J; we are going to do this pairwise element wise subtractions. We are going to see what is that element corresponding to the row I and column G in The Matrix a, which we say is a<sub>ij</sub>; we are going to subtract from this the element in the row I and column G that comes from Matrix B, and this will give us our new Matrix, which is A - B.

So let's now look into an example. In this Matrix, Matrix um uh a and Matrix B are used, and Matrix a is of the size 3x3 Matrix 3; Matrix B is of the size 3 by 3. In order to obtain A + B, what we are doing is that we are performing element wise additions. Now let's verify this. So what we are doing here is that we are saying A + B, let me actually get a larger area here. So let's say we have the two matrixes; I want to add the two in such way that we do everything one by one such that the idea of A + B and addition of the matrices will make sense. So we want to find out A + B; for that, what we are going to do is that we are going to make use of this definition that A + B, and then i<sub>j</sub> is equal to a<sub>ij</sub> + b<sub>ij</sub>, which is a fancy way or mathematical way or describing that for each element we need to go a look for the row I and column J and take that element from the um column from that uh Matrix a and from The Matrix B. So this means that for A + B, this is going to be a matrix that will have the same number of rows and the same number of columns as two matrices because both A and B are 3x3, which means also their sum is going to be 3x3. So this going to be 3x3, and here we are going to do—so we are going to take for the first row and the First Column, so for A + B<sub>11</sub>, so first row and First Column, we need to go to the first row and First Column of Matrix a and the first row and First Column of Matrix B, and we need to add this two elements. So we need to do 1 + 1, and then we need to go on to the second column, so the first row and the second column, which means that we need to be here in both matrices. So here we have 0 + 2, and then we got 2 + 3, and then we got 0 + 0. So you can see it in here, and then we have 1 + 0, and then we have 3 + 1, 0 + 1, and then 0 + 2, and then 1 + 3, which gives us—so, so 1 + 1 is = 2, 0 + 2 is = 2, and then 2 + 3 is = 5, 0 + 0 is = to 0, 0 + 1 is = 1, 1 + 0 is = to 1, 0 + 2 is = 2, and then 3 + 1 is equal to 4, 1 + 3 is equal to 4, which means that our A + B is equal to this Matrix that we got in here. So you can see that we are getting exactly what we uh what we have here, only we have done it manually one by one. So the same idea holds exactly when we have A - B, only instead of adding you will have to do here minuses, so minus, minus, so everywhere minus, so 1 - 1, 0 - 2, 2 - 3, etc.

Let's look into another addition. So in this case, by definition, it is defined as this element wise uh of the adding of these two matrices. Here the only difference in this definition is that it's saying it's calling this A + B as C. So this new Matrix that we are getting as a result of adding A to B, it's calling C. So basically it's the same as calling this Matrix as C. You will see also this type of definitions. So in this case, The Matrix C is equal to A + B, which basically means that for each row, row with index I and with each column with index J, go and look for row I and index J, take the corresponding elements from Matrix a and Matrix B, add them in order to get that corresponding element in our new Matrix C. And you can see that in this example that's exactly what we are doing. We have A, we have B; we are taking this element and this one, so 1 + 1, we are getting here two, and then 0 + 2, we are getting two here, 2 + 3 is 5, and then 0 + 0 is = to 0, 1 + 0 is = to 1, and then 3 + 1 is = to 4.

So now we already go to the next topic, which is about scalar multiplication of a matrix. So by definition, scalar multiplication of a matrix A by scalar Alpha results in new Matrix where each entry of A is multiply by Alpha. The idea of scalar multiplication matrices is actually quite similar to this idea of scaled multiplication in vectors. So uh we have already seen in the lecture of the vector multiplication that when we were having this scaler C and we had this Vector a, then uh when we are multiplying C, which is a real number, with Vector a, then we simply need to take all the elements of vector a, so A1, A2, or all the way down to a<sub>n</sub>, and we need to multiply them by this same scaler, so C. This is what we were doing with vectors, and that's exactly the idea behind matrices. And when uh during the scalar multiplication of matrices, only instead of multiplying only just one vector with this scaler C, now we need to apply this to all the rows and all the columns. So here we got this one column, and Matrix is simply a combination of multiple vectors, which means that we need to multiply all these elements of all the vectors of all the columns in this Matrix. So let's actually look into a specific example. So in this case, we have a matrix A, and this Matrix A is this thing, and we have a scaler which is three. So in here, our Alpha is equal to three, or you can call it C or anything. So you can see that when we are scaling The Matrix with a scaler, in this case three, with this Matrix, what we are doing is...

That we are simply taking each of these elements and multiplying it with the scaler. So, 1 by 3 is 3, 2 x 3 is 6, 3 x 3 is 9, and 4 x 3 is 12. This is the idea behind this entire scalar multiplication of a matrix.

In more general terms, if we, for instance, have a matrix A, so let's actually look into a high-level general example when we have the A matrix M by n. So we got M rows and N columns, and we want to get a scalar multiplication of this matrix and um scaler that we have here as in our definition. It is defined by this alpha. Alpha is just a number; you can go C, you can quote B anything. So, in this case, our scalar alpha, alpha is coming from R, so it's a real number. So, alpha times A is then simply equal to this new matrix where all of these elements are simply multiplied by this scale. So I would just take over all these values A1 up to AM1 and then A12, A22, all the way down to AM2, and then let me also add the last column just for fun here, A2N and then here AMN. So here, this new scaled mul multiplies, so scaled uh matrix A, so alpha * A is simply equal to alpha times all these elements are simply multiplied by this scale. It is as simple as that. So that's the simple idea behind uh matrix uh scaling. So when you are doing scalar multiplication of this matrix, you simply take all the values and you multiply them element by element per row and per column by that single scalar alpha. Do note that you are multiplying them all without exclusion with exactly the same number, which is that alpha.

Let's now look into the definition of matrix multiplication. So here we are no longer multiplying a matrix with a scaler, but we are multiplying a matrix with a matrix. So the product of an M by n matrix A and an N by P matrix B results in an M by P matrix C, where each entry Cij is computed as the dot product of the ith row of A and the jth column of B.

Now, what does this mean? Firstly, let's look and unpack this part of the definition. So we got matrix A that is M by n, and then we got matrix B which is n by P. What this means is that in this case matrix A has M rows and N columns, and matrix B has n rows and P columns. So this is then simply the dimension dimensions of the two matrices. So then it's saying that by definition the product of these two matrices, so the product of A and B, the product of the two matrices is equal to matrix C, and each entry Cij, so Cij is computed as the dot product of the ith row of A and the jth column of B. Now this part might seem a bit difficult, but once we look into the actual example and we illustrate this on our common high-level general expressions of matrices A and B and multiplication, this will make much more sense. For now, before coming to this one, I just wanted to refresh our memory on one thing I said before when discussing also this idea of improving uh this uh different properties of vectors that when we want to multiply a vector with a matrix or matrix with matrix or vector with a vector, we need to ensure that from the first element the number of columns is equal to the number number of rows of the second element. This is also very important for this specific case and just in general for matrix multiplication. So you can notice here that the number of columns here is equal to the number of rows in here, and the order is very important. So in case of matrix multiplication, the order is really important, which means that if you have a matrix A and you want to multiply it with a matrix B, then the number of columns of A should be equal to the number of rows of B; otherwise, you cannot multiply those two matrices with each other. So in case you got a matrix A that doesn't have the same number of columns as the rows of number of the matrix B, then there are some alternative things that you can do, including this idea of the transpose that we saw also doing when computing the dot product between this vector A and vector B. That's something that we also do in programming when we are dealing with this matrix and we want to compute this relationship between two matrices, but the number of columns of one of the first one is not equal to the number of rows of the second one. We are simply uh manipulating these matrices or moving some data if that's not hurting our problem, maybe uh flipping, so transposing our matrix or applying any other source of operation to it to ensure that the two matrices that we are multiplying with each other, the first one's number of columns is equal to the second one's number of rows. So that's just the low and that's something that you should follow if you want to multiply these two matrices. All right. So now let's move on onto the idea of multiplying and dot product. Let's look into a specific example, and this will um help us to understand this process better.

So before doing that, I just want to quickly show you this general idea. So if we have a matrix A that is M by n, which means that it looks something like this, like A11, A21 up to the point of Am1, and then here we got let's say A12, A22 up to the point of AM2, and then at the end we got AMN and here we got A1n. So let me also add this A12N and we got a matrix B. This matrix B is n by P, so it has n rows and P columns, so we are fine in terms of dimension here, and we got here B11, B21 up to the point of Bm, sorry, BN in this case, let's not confuse the letters, so Bn1, B12, B22 up to the point of Bn2 because n now is the number of rows for matrix B unlike for the matrix A, up to B1P and here B2P and here up to the point of B, and then NP. This is the last element. In order to perform um multiplication between these two matrices, so to obtain a matrix C which is equal to A * B, what we need to do is we simply need to take pair case, so pair row for the row I, for instance, we need to take this element, so this row, and we need to multiply it with this, so we need to find a dot product between this row and this column. Then we need to move on onto the next one, and then for the second element we will then take this row and we will multiply it with this one. So this is then something that we need to do in order to obtain these elements, and you might have already noticed that we got this M by n and n by P, so you might have already guessed what will be the dimension of the C. If we got that the dimension of A is equal to M by n and the dimension of B is equal to N by P, then the result matrix after multiplying the two matrices C will be will be having a number of rows equal to this and the number of columns equal to this. So this middle part basically disappears, and the number of rows of the first matrix will be then the number of rows of this result matrix C, and the number of columns of the second matrix, so matrix B, will then be our final number of columns. So we will then have a matrix C that will have a dimension, so dimension, so dimension of C will then be equal to M by P. So we will have M rows and P columns. So how we are going to compute this? So for Cij, which means row I and column J, let's look into the definition of it. It's saying Cij is computed as a dot product of the ith row and the jth column, so each row from A and jth column of B. What is the ith row of A? The ith row of A is somewhere here, so each row of A it is uh Ai and then I1, then Ai and then I2, and then Ai and then I3, dot dot dot, and then Ai and then we got in total n columns, and we always do the transpose right when computing this um dot product. So we then take the transpose, so we take this row, row I, and we multiply it, so we do the dot product between this one, this is the Ai and the Bj. This is column J; it is somewhere here, so it is B and then we got the first element which is one and then J, and then B2J, B3J, dot dot dot up to B and then in total we got n rows in B, so n and then the J is the column, so it's the same, so this is then the dot product between ith row that comes from matrix A and the jth column that comes from matrix B. So it's always like that actually, so we always take row by row, so we take this different, so every time we take just a row and we multiply with the corresponding column, and then we get the dot product between this row that comes from the first matrix and then the column that comes from the second matrix in that specific order in order to get our dot product and that specific volume. And what is this amount actually? So when we calculate this dot product, you can quickly see that we have Ai1 multiplied by B1J plus Ai2 multiplied by B2J and then dot dot dot AiN multiplied by BNJ, and this new matrix C will then have all these elements, so C11, C21 and then C31, dot dot dot and then see the last row as the number of rows of C is m, Cm, see here M, so C and then here it will be 1, 2, C22, C3 and then 2 up to the point of Cm and then 2, and then here the last col will be C1 and then P is the number of columns in C, so C1P and then C2P and then C3P, that dot dot dot and then Cm and then P. Okay, so this is what we get; this is our final matrix C when multiplying matrix A and matrix B. So let me clean this up. C is to now if you want to find out what is C11, you can easily fill in this general formula that uh that we just calculated, the I is equal to 1 and then J is equal to 1, and this will give you C11 by using this formula. If you want to get the CMP, then just fill in the I is equal to M and then J is equal to P in order to get this value Cmp. So you can already see the amount of calculations you need to do in order to get all these elements from these large matrices A and B.

Let's actually look into a simple example to clarify this. So we have a matrix A here and matrix B here, and we want to do a multiplication of the two, and we have just learned how to do it. Let's actually do it one by one. So we got a matrix A which is equal to 1 2 3 4 with dimensions 2 by 2, then we got a matrix B which has values 2 0 and then 1 2, so it is 2 x 2, and I want to find what is C that is equal to A * B, and I know already by looking at these dimensions that C is going to be equal to 2 by 2. So you might recall that I said that when looking at this final result, the number of rows of the final um matrix will be this, so the number of rows of the initial matrix A, and then the number of columns of this final vector C will be the number of vectors, number of columns of this second matrix B, so two. Therefore, I know already before even doing calculations that the uh product matrix C equal to A * B is going to have a dimension 2 x 2. Let's actually do a calculation to check this. So C is then equal to A * B and it's equal to 1 2 3 4 multiplied by 2 1 0 2. Okay, so I expect to have four different elements here, here, here, and here. So to obtain the C11, so it is C11 in here, what I need to do is that I need to look at the first row and the first column in here, here, so first row from A and the first column of B, and I'm doing the dot product, which means 1 * 2 + 2 * 1. 1 * 2 is 2, 2 * 1 is 1, so here I'm getting 1 * 2 + 2 * 1, which basically gives me 2 + 2 and that's equal to 4. So here I'm just writing down 1 * 2 + 2 * 1. Now when I want to get this value, which is C12, this means that I want to get the first row and the second column, and that's exactly what I'm doing, so I'm going back and I'm saying let's look at the first row, but this time we will look at the second column from the from the matrix B, so 1 * 0 + 2 * 2, and then I do the same, only this time for the second row, which means I'm picking this row and then this column, so it is 3 * 2 + 4 * 1, and for the final element C22 I'm taking the second row and the second column, which gives me 3 * 0 + 4 * 2. Now what does this give me? This gives me this 4 x 4 matrix where 1 * 2 + 2 * 1 is 4, 1 * 0 + 2 * 2 is 4, 3 * 2 + 4 is = to 6 + 4 which is 10, and then 3 * 0 + 4 * 2 is = 8. So let's check 4 4 10 8. That's is exactly what we have here. So as you could see here, the idea is that every time to follow what element I'm looking for for the Cij, and then I just go to the ith rows from the first matrix and the jth column from the second matrix, and I do the dot product of the A and then I and then K, let's say, so I'm going to the ith row from the first matrix and I'm taking all the elements, which means I don't even need to mention this index, it just means the entire ith row coming from the matrix A, and then I'm doing the dot product between this row and the column that comes from the matrix B, which means B and then J, which then will give me the Cij. So I'm looking at this and taking this, multiplying this dot product, and this gives me the first element, then the first row and then the second column, which gives me the uh second element in the first row in my matrix, so this one and so on. So hope this makes sense. Uh if it doesn't make sure to reach out because it's a very important concept, and uh let's also look into another example to make sure that we got this right. So in this case, as you can see, we have another matrices, so set of A and B matrices again 2 x 2, a simple one, and we want to know what is A, so let's say we call the C, we already know C should be 2 x 2, and what we are doing is basically for C11 we are saying let's look at the first row, so first row and the first column coming from the second matrix B, and let's do the dot product, so 2 * 1, 2 * 1, 2 * 1, 4 * 5, 4 * 5. We get this, and then when we want to find what is C, oh what is C, and then 1 2, so in the first row but in the second element in our final matrix, so I is equal to 1 and J is equal to 2, it means we need to look at the first row from the matrix A, but this time the second column from the matrix B, so it is 2 by 3, 2 by 3, 4 * 7, 4 * 7, and this gives us a number 34. Even if you calculate, you can see that 2 * 1 is equal to 2, 4 * 5 is 5, so 2, 4 * 5 is 20, so 2 + 20 is 22 in here, and then you can do the rest of calculations, and this will be a good practice to see how we can do a basic matrix multiplication. The idea is actually quite straightforward when it comes to multiplying it; it just it comes with a practice when we see all this uh much bigger matrices. So um this is another example; I will leave this one to you to complete it, just uh to keep in mind we always do uh so we always look at the dimension first, and here 2 x 2 and 2 x 2 which gives me an impression already what I can expect the result will be 2 x 2, and when it comes to the uh cross elements, just ensure to always look to the ith row and the jth column, this comes from matrix A and this column from matrix B, take them, compute the dot product, and then you will find your C, your final result. Let's call it um Kij because in this case we have a matrix C already.

Welcome to the module 4 of this course when we are talking about matrices and linear systems. So in this module we are going to dive deeper into this uh idea of linear systems with matrices and solving linear systems using different techniques, and specifically we are going to learn the uh concept behind solving linear systems using matrices named Gaussian elimination and Gaussian reduction. Welcome to the module 4 of this course when we are talking about matrices and linear systems. So in this module we are going to dive deeper into this uh idea of linear systems with matrices and solving linear systems using different techniques, and specifically we are going to learn the uh concept behind solving linear systems using matrices named Gaussian elimination and Gaussian reduction. So uh we are going to solve linear systems um with M equations and N unknowns using this idea of our commented coefficient matrix and then reduce row echelon forms, and we are going to see many examples. We are going to solve these problems, and we are going to do it step by step such that this very important concept from linear algebra will be uh entirely demystified for us. So we are going to learn each of these concepts one by one, and we are going to uh solve different problems like before with examples such that at the end of this module we will be clear on how we can use this idea of Gaussian elimination and Gaussian reduction in order to solve a linear system with matrices. So uh solving linear systems is crucial in linear algebra for many reasons. It helps to um provide this fundamental method for understanding and working with linear equations, which are the uh fundamental and backbone of various mathematical but also um the applied sciences like uh statistics, data science, machine learning, deep learning, and engineer artificial intelligence. So uh here is a high-level overview of its importance and applications. So uh it's a basic but essential part of linear algebra. It introduces key concepts such as matrices, vectors, determinant. So we need to know um and we need to incorporate all these different concepts that we already learned as part of previous modules in here because we need to know um the idea of matrices, the dimension of the matrices, uh this idea of coefficients versus unknown, so variables in our model, homogeneous versus non-homogeneous, how we do the indexing, so coefficient labeling in our matrices that we learned before. We also need to know this uh multiplication between a matrix and a vector, multiplication between matrix and matrix, the idea of transpose and all these different topics that we saw before uh coming to this specific module. So the uh solving linear systems has a wide range of applications when it comes to the um usage of it beyond linear algebra because it's being used also as part of many other linear algebra concepts including the um dimensionality reduction techniques, uh the decomposition techniques like um eigenvalues, eigenvectors, uh all these different mathematical and linear algebra concepts itself, they rely on this idea of solving linear systems with matrices, but we also go beyond the application mathematics when it comes to solving linear systems.

It's used in physics. It's used in engineering, um, and just in general, also even in finance, economics, uh, in different uh investment strategies when we are talking about quantitative finance and using data. Linear algebra is used in order to understand what kind of investment decisions we can make. So, in order to set some uh constraints or objectives, and in the essence, solving linear systems is kind of this getaway to both understanding deeper mathematical theories but also applying these concepts to solve real-world problems across various disciplines and domains, including data analytics, data science, machine learning, deep learning, artificial intelligence, and much more.

So, without further ado, let's get started. So we will start with the refreshment of this general linear system that we saw when we were beginning this uh unit. So we saw that we could represent a linear system with m equations and n unknowns with this format, where we said that we had this um m equations. So you can see we got in total m rows, and we got n unknowns which were referring to the columns. So then we also saw this idea of indexing a<sub>ij</sub>, where i was the index corresponding to the row and then j was the index corresponding to the column. And we said that these coefficients a<sub>11</sub>, a<sub>12</sub> up to a<sub>1n</sub>, and then the same holds for all the other equations or the rows, these were ways for us to understand where exactly we are in our system, what is the variable corresponding to which we are dealing, that coefficient, and what that coefficient will basically tell us: in which equation we are and corresponding to which variable or unknown we um have the coefficient. Because this a<sub>11</sub> holds entirely different information that is a<sub>23</sub>, and then this a<sub>mn</sub>, this a<sub>mn</sub> tells us in the m equation what kind of information we know and the relationship between this when it comes to the n variable or n unknown.

So this a<sub>11</sub>, a<sub>12</sub>, all these a<sub>ij</sub>'s, those are known because we are saying we can estimate them; those are constants, they come from ℝ, whereas the x<sub>1</sub>, x<sub>2</sub> up to x<sub>n</sub>, those are variables; those are things that we assume that either we don't know or in the future we will see that those are basically the features that we got in our data. And then this b<sub>1</sub> and b<sub>2</sub> up to b<sub>m</sub>, so far we have been kind of ignoring them. We already touched upon this when we were differentiating this idea of non-homogeneity and homogeneity of our system, but in this specific module we are going to finally make this b in use; we are going to make use of them, and we are going to see how we can transform this general linear system into uh concepts like matrices and vectors that we already have seen at the end of the previous module. So we are going to bring all these different topics together to solve this general linear system, and very uh soon it will also become much more clear why we want to do that, because solving linear systems it means that we are um getting an understanding uh what what is the exact relationship between these coefficients and the unknowns.

Now keep this notation in mind where we have all these different equations; you notice here the pluses. So we have the common notation of equations that we know from high school and pre-algebra, because we are going to transform these linear systems into matrices and vectors and the products of them. So we are going to go from this uh simpler notation to a bit more fancy notation, and we are going to bring our linear algebra into use in order to solve a system of m equations with n unknowns. So let's go to that bit more advanced notation. So here we go. So this is the matrix representation of linear systems, and a linear system of equations can be represented compactly using these matrices. So we already have seen this matrix; this is very familiar to us; we have been using this a lot, and this is the m by n matrix A. So this is the matrix A. You can note here we have a<sub>11</sub>, a<sub>12</sub> up to a<sub>1n</sub>, and then we uh we have all these different rows and the columns. So another thing that you can quickly notice is that we are dealing with a similar dimension for matrix A. So we have m by n; you can see we have m rows and we have n columns. Something that you can note when you are looking at the last column and the last row, and then uh keeping in mind that we have always a<sub>ij</sub>, where the i is the row index and then j is the column index. So that gives us understanding what is the dimension of A, which is m by n, and then we have now a new component here, which is something new here, which is this vector. So we got a column vector here x<sub>1</sub>, x<sub>2</sub> up to x<sub>n</sub>. So we already see here from the last element that most likely we are dealing with an n by 1 vector. And how I can know this and how I can be assured by this is because first of all the index is n, and secondly we already have learned: for us to be able to multiply a matrix with another matrix or matrix with another vector, the middle elements of the two, so those things they need to be the same. So we have learned that the number of columns of the first multiplier, in this case the matrix A, needs to be equal to the number of rows of our second multiplier, in this case the number of rows of a vector X. So you can see it in here. So here we got n columns as part of matrix A; here we got n rows as part of vector X, which does make sense, which means that if we are writing like this and we are multiplying the two, then we are indeed dealing with the case when those two should be the same. So we are definitely safe to assume that our vector X it needs to have n rows and only one column. So of course the one could have been different, but we have already seen another thing before. So why do we know that this is right? We are indeed dealing with a single column vector. Well, if we go back in here, we are simply taking this system and we are trying to rewrite it in this form, and what have we seen here that we only got n different unknowns, something that was also given, which were x<sub>1</sub>, x<sub>2</sub> up to x<sub>n</sub>, which we can also rewrite in this format: x<sub>1</sub>, x<sub>2</sub> up to x<sub>n</sub>. So basically a column vector; that's exactly what we have done here. So we have taken all the x's and we have put it in the vector, and we are saying let's multiply this matrix with this column vector, and we will get this. And how we can know this? Well, let's first look into this first part, and then we will understand why we have here this um vector of b. Let me clean this up. Well, what we have done simply here is that we have rewritten this part, this part as a multiplication of a matrix and a vector. So we have said let's take this; we have already learned how we can perform the multiplication between a vector and a matrix; we have seen that and how we can do that. Now we are basically doing the opposite. So we have already the final version; we have all these equations. You can see that we have taken a<sub>11</sub>, we have multiplied with x<sub>1</sub>, a<sub>12</sub> we have multiplied with x<sub>2</sub>, and a<sub>1n</sub> we have multiplied with x<sub>n</sub>. So basically what we have done here is that we have taken the a<sub>11</sub>, a<sub>12</sub>, and then a<sub>13</sub> ... and then a<sub>1n</sub>; we said let's take this and then multiply with x<sub>1</sub>, x<sub>2</sub> ... x<sub>n</sub>, which is basically the dot product of this row vector with this column vector. So only the first equation can be rewritten as a dot product between this vector that we see in here and this vector. All right, perfect. But you can say, well, we just got uh we not only got one equation, but we got m equations. This already gives us an idea what we have done in the left-hand side; what I just did for one equation we can also, of course, do for multiple equations. So basically what we have done was that we have set for equation one, or I can also say for i = 1, we have taken a<sub>11</sub>, a<sub>12</sub> ... a<sub>1n</sub>, because we got n columns or n unknowns, and we multiply this with x<sub>1</sub>, x<sub>2</sub> ... x<sub>n</sub>, and this dot product, this is basically A<sub>1</sub><sup>T</sup> and then X, if we call this X and then this. If you open this up and if you calculate this, you will see very quickly that you get a<sub>11</sub>x<sub>1</sub> + a<sub>12</sub>x<sub>2</sub> ... a<sub>1n</sub>x<sub>n</sub>. This is simply the the proof how we can go from here to here. Now if we do, if we show the same in the same way, you can actually go ahead and prove for yourself that i = 1, we simply have this vector a<sub>11</sub>, a<sub>12</sub> ... a<sub>1n</sub> multiplied by x<sub>1</sub>, x<sub>2</sub> ... x<sub>n</sub>, but we also have the same for i = 2, only with a different row vector, which is of course a<sub>21</sub>, a<sub>22</sub> ... a<sub>2n</sub> and then dot product with x<sub>1</sub>, x<sub>2</sub> ... x<sub>n</sub>. So basically the column vector X doesn't change, but we need to change the coefficient that comes because uh in each equation we have a different set of coefficients corresponding to this X, and not you can see that here coefficients are different; their i or the row index is changing because here we have i = 1, then i = 2 ... i = 1, i = m. In the same way here we are also doing the same. So here you can see we have exactly the same story, only now a = m for our last row, which is a<sub>m1</sub> and then a<sub>m2</sub> ... and then a<sub>mn</sub> and then multiply. So the dot product with vector x<sub>1</sub>, x<sub>2</sub> ... x<sub>n</sub>. Given that for all these cases we got the same column vector, you can easily see see how we, instead of writing down this equations as separate row vectors, we have seen also the same doing previously, and instead we will just create one coefficient matrix. So we can say instead of just doing this separately, we're also safe to do it um in a combined way. So therefore we are rewriting this all, and we are saying we can rewrite all these equations now as a<sub>11</sub>, a<sub>12</sub> ... a<sub>1n</sub>, so as I had before. So I'm taking just this part, and I'm adding here my column vector that I had to multiply this um A<sub>1</sub><sup>T</sup>x<sub>1</sub>, x<sub>2</sub> ... x<sub>n</sub>. So basically I'm just taking this that we had and I'm writing over in here. So this is my first row, but then of course I can also add; there is nothing that I'm changing if I'm adding a similar story, only this time I'm using this second row vector a<sub>21</sub>, a<sub>22</sub> up to a<sub>2n</sub>, because at the end of the day I'm doing similar multiplication with x in here too. So a<sub>21</sub> and then a<sub>22</sub> ... a<sub>2n</sub>, and then of course the same holds for a<sub>m1</sub>, a<sub>m2</sub> for my last row and then up to a<sub>mn</sub>. So if you open this up, you will quickly see that what you will get is that for the first equation I will get entirely separate amount, so a<sub>11</sub>x<sub>1</sub> + a<sub>12</sub>x<sub>2</sub> ... + a<sub>1n</sub>x<sub>n</sub>. So I'm simply computing the dot product between this row vector and this column vector. We have already seen how we can do a calculation like this; this should seem very familiar now. I'm only doing it at a scale; I'm doing I'm adding here also the second row, so a<sub>21</sub>, a<sub>22</sub> ... and then a<sub>2n</sub>, and then here of course I forgot to add the x's, so a<sub>21</sub> * x<sub>1</sub> + a<sub>22</sub> * x<sub>2</sub> + ... and then a<sub>2n</sub> and then x<sub>n</sub>, and then I'm doing this all the way down to the last row, a<sub>m1</sub>x<sub>1</sub> + a<sub>m2</sub>x<sub>2</sub> + ... and then a<sub>mn</sub>x<sub>n</sub>, and then what I'm doing is that I'm saying this is equal to this. You are free to check it yourself; um, you can see that once you write down even all the smaller elements in here, you will see that you are getting exactly the same. This is simply uh another way of writing that you will take all these different um vectors A<sub>1</sub><sup>T</sup>, A<sub>2</sub><sup>T</sup> ... and then you also represent your last row by a vector, and then you take this, which is basically m by n, and then you do dot product with a vector x<sub>1</sub> and x<sub>2</sub> ... x<sub>n</sub>, and why am I allowed to do so? Because I have m different equations, so equation one, equation two ... equation three, those are the coefficients of those corresponding equations, and those are the unknowns, and given that the unknowns are exactly the same across all the equations, so as you can see here the unknowns they don't change, so this means that every time I need to multiply, compute the dot product between a different coefficient, so a<sub>11</sub>, a<sub>12</sub>, and then it changes to a<sub>21</sub>, a<sub>23</sub> ... a<sub>2n</sub>, but the unknowns they do not change, so the vector, the unknown vector that I need to to multiply them, it doesn't change every time; it's this. So in here I have the same, in here I have the same, and in here I have the same column vector. This means that I can simply rewrite the set of dot products into a larger uh system where I have this coefficient matrix, which is A, so and we are calling it coefficient matrix; we saw already before why, because those are the coefficients corresponding to the unknowns; we are storing them in one place, and then we are storing the unknowns x<sub>1</sub> to x<sub>n</sub> in a column vector because we got n unknowns, and then the final piece of this puzzle is what we see in here in the right-hand side after the equation sign, of course; I'm talking about the b. So we got here b<sub>1</sub>, b<sub>2</sub> up to b<sub>n</sub>. You might have already guessed that we are uh dealing with um a vector, not a matrix, because we got m equations and m b's, which means that we can represent this as b<sub>1</sub>, b<sub>2</sub> ... b<sub>m</sub> as a column vector with m rows because we got here one row, two row ... here m row, and then by given that it's just a one column here, one. So basically we are saying for equation one, a<sub>11</sub>, a<sub>12</sub> ... a<sub>1n</sub> multiply it, so dot product with X, so x<sub>1</sub>, x<sub>2</sub> up to x<sub>n</sub>, this should be equal to b<sub>1</sub>. This is exactly what we have; let me clean this up. This is exactly what we have here: we say a<sub>11</sub>x<sub>1</sub> + a<sub>12</sub>x<sub>2</sub> ... a<sub>1n</sub>x<sub>n</sub> is equal to b<sub>1</sub>; that's what we have here because we are saying the dot product of this row vector that we have just used to rewrite this thing, that product of this uh coefficient with the corresponding unknown is equal to b<sub>1</sub>, and we have already seen that it is simply this thing that we are rewriting this first equation we have is equal to b<sub>1</sub>, as you can see in here. Therefore, when we do exactly the same, it's simply equal to here b<sub>2</sub> for the second equation ... and then here is equal to b<sub>m</sub>. Given that from the left-hand side here we already have seen that this is simply equal to coefficient matrix A and then dot product with X, where X is this vector, it means that we can also take over this equal sign from here, and the only thing that is left from us is to represent this column vector in here, which we can call a vector B, and that's exactly what is in here. So soon we'll also see the clear definition of it, but for now let's see how we uh we have used this linear system in order to transform it in the notation with matrices.

So we have taken this coefficients, and we have seen that in here we can represent these coefficients as a vector, a row vector, so a<sub>11</sub>, a<sub>12</sub>, a<sub>1n</sub> multiplied with a column vector x<sub>1</sub>, x<sub>2</sub> up to x<sub>n</sub>. So this is exactly the first equation, the left-hand side. So we say that this is the first dot product, this is the second dot product up to the last row, the last dot product. So we have already represented this dot products as um multiplication, so product of a coefficient matrix A by column vector containing all these unknowns x<sub>1</sub>, x<sub>2</sub> up to x<sub>n</sub>, X, and we said let's take over this equal sign, and given that the same logic row-wise holds also in here, so we say this is row one, this is row two ... row m, and everything that holds here for the rows holds here as well; we need to take over everything in a consistent way. Therefore we are saying let's also represent this in terms of the vector; if we are doing that in here, therefore we are calling this a b, and this b's are just real numbers b<sub>1</sub>, b<sub>2</sub> up to b<sub>m</sub>, and this is how we can transform and we can rewrite this system with m equations and n unknowns into a product in terms of matrices, or coefficient matrix A, a column vector containing the n unknowns, and then a vector, a column vector B that contains all these different values b<sub>1</sub>, b<sub>2</sub> up to b<sub>m</sub>. So this basically rewriting what we have in here. Hope this makes sense, but if it doesn't, we are going to see this again in the definition; I just wanted to warm up with this before we move on onto the uh official definition.

So a linear system of equations can be represented compactly using matrices, and we we just saw that an m by n matrix A multiplied by an n vector X results in an m vector B. So you can see it in here. So this is an m by n matrix A, this is the n by 1 column vector X, and this is the m by 1 column vector B. And one thing that we can see very quickly is that what we are doing here is that we are transforming that uh equations into a matrix representation, but nothing changes with the equations; we can still write them down like that; it's just by using matrices and vectors we will soon see how we can use this in order to solve this problem. So why are we doing this? Uh, when we have just two equations with two unknowns, we know from high school that we can quickly solve that; we can just write down for instance x<sub>1</sub> + 2x<sub>2</sub> = 3 and then 3x<sub>1</sub> + 4x<sub>2</sub> = 6; we can quickly solve this problem by uh, you know, writing down the x<sub>1</sub> in terms of 3 and 2x<sub>2</sub>, filling that in in here in the x<sub>1</sub> and then solving the problem because we got two equations with two unknowns. The problem becomes much more complicated when we got three equations with three unknowns, five equations with five unknowns, 100 equations with 100 unknowns. Therefore, to do this uh solving of the system of equations at the scale, we need this matrix notation, and that's exactly what we are doing in here. So uh let's also quickly check the dimensions such that we confirm that we are uh on the right path before moving on to the uh practical material. So here we got matrix A containing the coefficients uh from our linear system, and it is, let me remove this, it is m by n. Then we got our X, so this thing is a column vector containing the uh unknowns in our system, and it contains n rows, so the X has a dimension of n by 1; it's a single column vector, and we are saying this is equal to, and that's just given; that's how we transform our linear system into this notation; this is equal to a column vector B, and the dimension of it is m by 1. Now let's check whether this makes sense. We know that already, and we just said that this two needs to be equal; it's fine. So we have the uh this uh criteria satisfied: so the number of columns of matrix A, coefficient matrix A, is equal to n, and then the number of rows of column vector x with unknowns is equal to n. So we can indeed do the multiplication. So we are correct in that notation, but uh is it true to claim that indeed here we got m different rows? Well, we have learned that if we take an m by n matrix and multiply it by an n by 1 vector or matrix, then what we are getting is that we first take over this, so the number of rows, this will become the new number of rows of our final matrix, and then m by and then this last

element so the number of columns of our last item so in this case DX, so it's equal to 1. This is our final number of columns, so the new dimension then becomes M by 1. And what do we see here? That indeed the result that we say we claim that this amount is equal to is M by 1, because if we are saying that the two amounts are equal to each other, then uh naturally they should have the same dimension. So A * X will have a dimension M by 1, which is coherent to what we are claiming that is equal to, which is M by one uh column vector B. All right, so now when we have um refreshed our memory on that dimensions, we are ready to move on onto the next slide.

So now we need to solve this linear system. So we need to find this set of unknowns that will help us to identify and clarify all the unknowns in our system. So one way, and the INF famous way of solving this linear system, is what we are referring to as Gaussian elimination. Gaussian elimination is the systematic method for solving linear systems. It basically transforms this system into what we are referring to as row echelon form, from which the solutions can be more easily identified. So in the next couple of slides and uh lessons, we are going to learn this idea of elimination. We will understand why is it actually called elimination, which is quite intuitive once we see the examples, and also we are going to learn this idea of row echelon form, reduced row echelon form, because those are very important concepts to understand this entire idea of solving linear systems using linear algebra.

So before moving on onto that, we need to know the basics, one of which is the idea of augmented coefficient matrix. So by definition, the augmented coefficient matrix of linear systems AX = B uh combines the coefficient matrix A and the vector B into one matrix. So you can see here the matrix A, and then we have here the straight line, and then B. This is basically this entire augmented coefficient matrix. It's a fancy way of saying, let's take all the unknowns from our system, which are the coefficients in our system, and then the constant uh in the B. So I have to say not nouns, but rather uh the uh values, so the scalars, coefficients, and let's take the coefficients um in our system which are represented in our coefficient matrix, and let's combine this with the information that we got in our system, which is B1, B2 up to BM, and let's combine the two and we call it augmented coefficient matrix. So we are simply taking the coefficient matrix A, you can see a11, a12 up to a1n, and then also in the last row am1, am2, amn. This should look very familiar because it's just simply this matrix, matrix A. We are taking it over in here, and then we are saying let's draw a straight line and then add here the uh column vector, our result vector, which is B. So we saw in here that is this thing. So let's take this part and this part, put a straight line in between, and this will become our new matrix, Matx.

So as you might have already guessed, in the A, in the matrix A, in the coefficient, we got M rows and N columns, which means that our new matrix will have at least M rows and at least N columns. Now, given that we are adding yet another column at the end, it means that the number of columns will increase, but one thing that is important to note is that the number of rows will not increase because both the coefficient matrix and the column vector B, they have the same amount of rows; they both got uh just M rows. So let me actually rewrite this. So we know that A is M by N, so it has M rows and N columns, and we know that B is M by 1, which means that it got M rows and one column. This means that if we combine horizontally, so if we put them together next to each other, the coefficient matrix A and B, this means that we will end up with the same rows, so we still have M rows, but now we will have N + 1. So here are the N columns from N from uh our coefficient matrix A, and here yet another extra one column that comes from the uh vector B. Therefore, we get the augmented coefficient matrix that has a dimension M by N + 1. So this matrix is key when we are applying these different row operations to solve our system of M equations and N unknowns. So we are basically bringing all these values in order to solve our problem and find the unknowns.

Now, before moving on onto actual uh definition of row echelon form, RREF, I wanted to provide here an example, so how we go from uh from this system of equations into this uh new form A * X is equal to B, and then how we can represent our matrix into this idea of augmented coefficient matrix. So let's assume we have the following system of equations: So we got X1 + 2X2 + X3 + X4 is equal to 7, and then we got X1 + 2X2 + 2X3 - X4 is = to 12, and then 2X1 + 4X2 and then + 6X4 is = to 4. Let's now first just write this equations with three equations and then four unknowns into this system where we transform this into this matrices and the vectors. So let's write it in this format of AX = B that we saw before, and then we will also write this in the um augmented coefficient matrix. So given that I know that we have three equations and four unknowns, the first thing that I will do is that I will get my coefficient matrix A. So for A, what I need to do is that first I need to understand what is the dimension of it and how can I know the dimension? I know that it should be M by N from the theory where M is the number of equations, the number of rows, and N is the number of unknowns. I got three equations, four unknowns, so this is 1, 2, 3, and 4, X1, X2, X3, X4, those are my unknowns. This means that I already know exactly what the dimension of my coefficient matrix A should be, which should be 3 by 4. Let's now go in and search those values and filling in here to get our coefficient matrix. So here I see that X1 has a coefficient of one, so whenever you don't see here a value, it means we have one times that variable because one * a number is simply that number. So here I got one, and then uh for 2X2 I got two, for X3 I got one, and then for X4 I also got one. And do keep in mind that in the columns I have to have the coefficients corresponding to this unknown X1, X2, X3, X4. This will be kind of a guide mentally to check every time whether we are dealing with the right coefficients. So now we go on to our second equation or our second row, which means I got one in here, see here one, and then two, two, and then minus one, and then here for my third row, third equation, I got two for X1, I got four for my X2, I got nothing for my X3, which means that it is zero because 0 * a variable is equal to zero, and then finally for my X4 I got six.

So now I have my coefficient matrix. We can rewrite very quickly the column vector containing the unknowns. We already know what the dimension of it should be because we have four uh unknowns, so it should be 4 by 1, by one, and it is X1, X2, X3, and X4, and this is my X. And finally, I see what my B is here. So we got three equations, so we remember that the dimension should be M by 1, which is corresponding to what we also see here; it's coherent. We got three Bs for three equations, so therefore I'm saying my B is equal to, it is 3x1 column vector, and is equal to 7, 12, and 4. Now the final result is simply, let me rewrite this in the form that we just saw, so I can write this equations, the system of three equations with four unknowns as 1 2 1 1, 1 2 2 -1, 2 4 0 6 multiplied by X1, X2, X3, and X4 equal to 7, 12, and 4. This is in the AX = B matrix notation. So this is the first part. Now we know how we can rewrite the system of linear equations with uh M equations and N unknowns into this form of AX = B. Let's now try the next step, which is writing this um system, rewriting it into this augmented coefficient matrix. So I want to get this AB. This one is actually quite straightforward now when we have our A and the B, because what this augmented coefficient matrix represents, it's simply taking the matrix A and then adding the vector B next to it in order to get our final uh A + B matrix. So not A + B, but basically A and then next to it the B. So what is my A? It is this thing, so it is 1 1 2, and then 2 3, and then 1 2 0, 1 -1, and then 6. So now this is the A with three rows and then four columns, and then here I'm simply putting this straight line like in here and then putting the B in there, and the B was this, so 7, 12, and 4. So 7, 12, and 4. This is the B, which was 3x1, 3 by 1, so one column and then three rows, which means that this final matrix, the augmented matrix AB, it should have three rows and five columns. So it's really important to um distinguish the multiplication and this unique form which is called uh the augmented coefficient matrix. So here we are not multiplying or adding; we are simply taking the coefficient matrix that we just found, A, next to the vector B, and given that they both have the same rows, that's very straightforward and easy. So this is about the augmented coefficient matrix. Let's now move on to the next concept, which is the row echelon form, or RREF in short.

So matrix A is in row echelon form if all nonzero rows are above any rows of all zeros. So all nonzero rows are above any rows of all zeros. Each leading entry of a row is to the right of the leading entry of the row above it, and all entries in a column below a leading entry are zeros. The leading entry in each row is known as the pivot or corner, corner entry. Now this is a whole mouthful, and um the concept itself is not straightforward when it comes to this definition, but once we do the actual examples, all this concept will be clear, and when we solve these multiple problems, and we are going to solve like three till five problems, it will become easier to understand this concept of row echelon form, what we mean by pivot or corner entry or pivot uh uh vector, what we mean by uh having all these entries in a column below a leading entry zeros, or what we mean with this concept of all nonzero rows being above any rows of all zeros. So in terms of the words uh this might seem bit complicated, and it is, but when we look into the example, this concept will be demystified. So um once we have the RREF, which will be uh will uh come back in a bit and we will look into the example to clarify this definition, let's quickly look at the next step that we need to perform, which is the reduced row echelon form.

So here you can see that this reduced is what is basically added in front of the row echelon form. So this was the row echelon form that we will now uh be discussing, and then once the row echelon form is done, we basically do a reduced row echelon form. So that is the next step, and by definition a matrix is in reduced row echelon form if, in addition to the row echelon form, so RREF, we also have every leading entry in a row that is uh equal to one. Also we are calling that we have leading one, and each leading one is the only nonzero entry in its column. So RREF, or the reduced RREF form, it simplifies the system for easier solution derivation. So in terms of explaining or in terms of formulating the uh RREF doesn't really uh sound that convincing because we need to demystify all these different concepts. So my suggestion would be let's move actual to the actual practical problem, problem, and once we solve this step by step with all these details, all these different criteria, this one, two, and three, and then the four and the fifth criteria, they will be demystified and will be very clear to you.

Let's now solve another problem, a similar one to what we just solved, only uh this time we will go a bit faster through all these steps, just to ensure that our understanding of solving a linear system with this multiple equations and multiple unknowns using this idea of Gaussian elimination, reduction, and all these different steps up to the point of the reduced row echelon form, we can go through quickly, and we will ensure that uh we end up uh with the solution uh to the system using those techniques that we have just learned. So here is my system of three equations and three unknowns this time, and like we learned before, the first step is to transform this into this augmented matrix, but first of course I need to know what is my A and what is my B, because uh transforming this and writing this in the um augmented matrix form A and then B, it means I need to know what is my A and what is my B. So let's write the first that down. So my step one in this case is to write it in the augmented, so augmented matrix, and for that I need to know what is my A. So my A is equal to 1 1 1, and then 2 3 1, and then 1 -1 and 2. This is my A, which I take from all these coefficients that I see in here, and then my B I can see in here from all this values, so it is 6, 14, and 8. This means that my augmented matrix is A and then B, and it's equal to 1 1 1, and then 2 3 1, and then 1 -1 and then 2, then straight line, and then 6, 14, and 8. This is my augmented matrix. This is my first step. So let me go ahead and remove this details because then I can take this over and I can then simply continue. So my augmented matrix is then 1 1 1, and then 2 3, and then 1 1 -1, and then 2, and then I have the values for my B which is 6, 14, and 8. All right, first step is done. Now we are ready to go on to the next step, step two, which is to apply all sorts of operations, and ideally by every time normalizing each row, so every time ensuring that here the leading entries are ones, we call it normalization, and then by elimination. So we saw that every time we were eliminating some of the variables to ensure that we will end up with the zeros in the uh lower diagonal of our augmented matrix, and we will have only the leading ones in our um only a single one and a leading pivot value, so the pivot entry in our augmented matrix per column, it will be the only element that will be uh nonzero, and the rest should be all zeros. So here I'm talking about performing all these different operations to ensure that we uh create and we come to this point of the row echelon form, and then from there we will go to the reduced row echelon form, so the RREF that we saw before. So let's do that. So the step two will be to uh get the RREF of A, or rather I have to say RREF, maybe we can combine the two uh and do it at the same time. So for that what I will do is that like before I would just draw this line because this will help me to keep track of all the operations that I'm performing later on to uh to check. One thing that we can notice here that the first row is already normalized, which means that I have a one in here, so I'm happy with that. I won't do anything to it. So uh what I need to do is to move on on t uh on the elimination of this um X1 from the row two and row three. So I need to find a way to ensure that this entries become zero, because I want to get here 1 0 0. So I want to have a single pivot entry in here, and the remaining values in this column under this one should become zero. So how can I do that? The first thing that I can do to eliminate this X1 from row two is subtract from this two times the row one. So you can see that if I take this row two, so this is my row two, this is my row one, and oh this is my row two, this is my row one, and this is my row three. If I take row two, this is my step 1, row two is equal to, and then I take the row two and I subtract from this 2 * row one. Let me remove this to open up some space. So if I take this, this row and subtract two times this one, then what I will get is 2 - 2 * 1, which is 0, 3 - 2 * 1, so 3 - 2 is 1, and then 1 - 2 * 1, which is -1, and then when it comes to the value of B, 14 - 2 * 6, so 14 - 12 is 2. So this is what I'm getting, and then given that I have already a normalized first row, so I have already one in here, I will just take this over for now, and then for the last column, for the R3, in order to normalize that one and in order to get R of this one from here, so first thing I want to do is to eliminate the X1 from here, which means I need to get rid of this one, and how can I do that? I can simply take the row three, so R3, and I can subtract from the R3 the R1, because both have one in here, and this one I can simply take this and then remove and then subtract from this the first row. I can uh then end up with basically a zero. So here if I take this one, I got this one, and then I subtract from this 1 * 1, so 1 - 1 is 0, -1 - 1 is -2, 2 - 1 is 1, and then 8 - 6 is 2. This is what I get after performing this eliminations, hence also the name um Gaussian elimination. So using this eliminations we are then getting rid of certain elements and certain variables. So this is what we get after performing these row operations, and now we are ready to go on to the next stage. So we see now that the um in this specific case the row two is already normalized, which means that here we already have a one, so we are happy with that; we want it to be like that. So I already see another coefficient that I want to get rid of, that I want to eliminate from uh X2, which is this one. So I want to get rid of this -2 and I want it to become a zero, and one way that I can do that is to take this row three and then add 2 * row two to this row three. So why am I not taking row one? Because row one has a one in here, and if I take that I will then um basically um bring me back to the point where I have I'm having here an element because zero, and then uh adding something that uh is related to this first row or subtracting from it a multiple of this first row, it will always lead me here this zero becoming a nonzero value, that's something that I want to avoid. Therefore, I want to perform some sort of operation to this row three by using row two, because that's my only way to eliminate this -2 from here without hurting my zero here, and therefore I'm looking at this row that already contains the zero here, so that's great, and I'm going to use row two to eliminate this -2, and the way that I can do it, given that here I have a minus, so I have -2, I need

To add to this 2 * Row 2. So, in step number two, I will then say that my R3 is equal to R3 and then plus 2 * R2. So let's go ahead and see what that gives us. So 0 + 0 + 2 * 0 is 0, and then -2 + 2 * 1 is -2 + 2, which is zero. So I'm already getting rid of the minus 2, which is great; that's exactly what I wanted. And then 1 here, I have a one, so 1 + 2 * -1, so 1 + 2 * -1, that gives me 2 - 2, or 1 - 2, which is minus 1. So I'm getting here a minus 1. And then here, when it comes down to the B element—so I'm talking about this specific element—so 2 and then plus 2 * 2, 2 + 4 is equal to 6. So here I'm getting a six. Okay, perfect.

So let me take over the row number two; it's already normalized, so I already have a one in here: so 0, 1, and then -1, and here I got a 2. So I just simply take over; I'm not doing anything at this stage to that row number two. So, oh, now the question is whether there is something that I need to do to my Row one. Well, the question is yes, because in our definition we saw that we wanted to have um this element ideally being zero, because we need to have here a zero for this column to have a single element of one in the pivot entry, which is this one. Which means that I need to make somehow this element one that we see in the row one being equal to zero at this stage. Let's see how we can achieve that. So we got a one in here and we got a one in here. So one thing that we can do is to take Row one and subtract from that Row two. In that case, I won't do anything to my first entry, my pivot, my pivot entry for my first row; I won't do anything to it because 1 - 0 is 1. But then I will get rid of this one because 1 - 1 is equal to 0. So let's go ahead and do that. So R1 is = to R1 - R2. So in that way, if I do 1 - 0, I'm ending up with 1; 1 - 1 is = to 0, so I'm getting what I just wanted. And then 1 - -1 it is 2 because it becomes +1 + 1, and then 6 - 2 becomes 4. So this one becomes four. This is the end result after performing these operations in Step number two. Perfect.

So now, uh, let's see what else we got in here that we want to change. Let me make this line longer. All right. So there is something that I still want to do in order to normalize the last row, because here, instead of having one, I have minus one. So firstly, I need to multiply this row by one, simply by a scaler, in order to ensure that I can have one in here instead of just minus one. So in Step number three, I will do R3 is equal to -1 * R3. So I will get 0 0 1, given, and then 6 * -1 is -6. Now, what else we need to do uh in terms of the rest of the values in our uh augmented Matrix? So we got one; let me actually write it over such that we see what is going on. So I'm taking over this and then -1, and then here I got 2, then here we got 1, 0, 2, and then 4. All right, so this is our augmented Matrix, and there are few elements that we want to get rid of: it is this element and this one, because we want them to be equal to zero. And how we can do that? So we got two things to eliminate: we want to eliminate from the row one this two, and we want to eliminate from this row uh to this minus one. How can we do that? Well, for this Row one, we can already see that we got here this one from row three, so it's super helpful in that aspect. So we can use that in order to get rid of that R2. So let me actually write it down in the uh step four. So step number four, what we can do is to take R1 and it will be equal to R1 - 2 * R3, because here, here I got one, and I can use this one in order to get rid of this two. So 2 - 2 * 1 is equal to uh 2 - 2 * 1 is equal to 0. So I can get a zero in here, and that's exactly what we want to do: we want to eliminate this x3e. So we want to eliminate this from here, this two from this first row. That's something that we can achieve by this. And another thing that I can also see—let me free up some space in here to write down these equations—so remove this, given that the steps are already written, you can always replicate that. So then the new augmented Matrix a will become: so here 0 0 and then one; I'm leaving the last R; it is already in a uh format that is desired, so I will leave that there. And then I'm going to update my first row. So first row R1 is equal to R1 - 2 * R3, which means it's equal to 1 - 2 * 0 is 1; 0 - 2 * 0 is 0; 2 - 2 * 1 is equal to 0; and then 4 - 2 * -6, so 4 - 2 * -6 is equal to 4 and then + 12, which is 16. So I'm getting a 16 here. And then for the second row, I also need to do elimination to get R—this time, me remove this—this time I want to get rid of this minus one, because then I will have um criteria satisfied that this should also be zero, because this should be zero for this pivot to be the only nonzero element in this column three. Only then I can say that I'm dealing with RREF, so that I have my Matrix in the reduced row echelon form. So for that, what I'm going to do is I'm going to make use of my third row again. So I'm going to take row two, and I'm going to subtract from row two—I'm actually going to add Row three to it—because here I have minus one, and if I add -1 to 1, I will just get a zero. So for that, I'm going to write that uh for the elimination. So for eliminating x3 from the second row, I'm going to write that R R2 is equal to R2 + R3. So then I'm going to get here 0 + 0 is 0; 1 + 0 is 1; and then -1 + 1 is zero. Now, why I have used actually the third row, not the first one, because I want to ensure that I want to do anything to my zero in here. If I do something with this first row, then I cannot do that; I cannot reach that; it's just counterintuitive, because I want to keep my work intact. I have worked uh in such a way that I can eliminate this X1 from the second row, and that's something that I don't want to go back to. Instead, I want to use something that is in the column three to get rid of this column, and that's something that we could do by using this element. So we always look at the shortest and easiest path to get rid of a coefficient in our Matrix, in our augmented Matrix.

Now let's see what is left. So R2 is equal to R2 + R3, which means that we need to add this two to this -6, and 2 + -6 is equal to -4. -4, there we go. So after performing all the steps, we end up with the following augmented Matrix. And if you can notice here, we not just achieved the row echelon form, but we have actually achieved the reduced row echelon form of our a. This is that form because we have on the entry. So our pivots are the ones in here, and we can see that all the criteria are satisfied. So here, on the lower diagonal—so the lower part of our Matrix—are zeros, and on the top of that also all the elements in this columns, they are all zero except of the pivot values, and they are all equal to one. And that's exactly what we wanted to have in order to say that we are dealing with the reduced row echelon form. Perfect.

So now we are ready to move on to the last final step to get this solution. And unlike the previous system, which was bit more complicated and it had multiple infinite solutions—so you might recall this example where we had to describe the final solution using this X2 and X3, because depending on different values of X2 and X3, which could take infinite values, we then could get a different um solution for our x uh Vector—so we we saw that when, for instance, the X1 is equal to 1 and X4 is equal to 3, then our solution was this, but then of course if the uh x2 and X4 were different, we would have gotten another solution, and the result of the linear system had infinite number of solutions. Unlike that one, this one, this example is much more convenient, because even from the RREF of this—so row reduced row form of the uh Matrix a—we can already see what is the solution of this problem, of this linear system. We had three equations and three unknowns, so we actually got to single solution; that's something that I can already see from here. And let's uh break it down to see how um I see that solution. So we transformed this to a system of equations; so we are basically doing it backwards. And now, after getting this RREF of Matrix a, we are going going back, and we are going to rewrite the system with the unknowns X1, X2, and X3. So here I see one, which means that the coefficient of X1 is 1, so 1 X1, which I will just write X1, X1. So I get X1 + 0 * X2 + 0 * X3 is equal to 16. And then 0 * X1 + 1 * X2 + 0 * X3—I'm writing down this zero such that it will make sense what we are doing, and in case next time you have a different system, it will be easier to follow—so 0 * X1 + 0 * X2 + 0 + 1 * X3 is = to -6. Now what is this? We get rid of all these zeros, so you can already see what is going on. We are ending up with X1 is equal to 16, X2 is = to -4, and X3 is equal to -6. There we go. So we got our solution for X, which is X1, X2, X3 is equal to 16, -4, -6. This is the unique solution to our problem. And if we clean this up and we write down in a nicer term—so I'm going to remove this—then we can say that the solution for our problem—so this problem that we saw in here—so the X, which is equal to X1, X2, X3, it's equal to 16, -4, -6, which you can see it also here. So one trick is always to look into in here: if you have the identity Matrix part of your augmented Matrix, then you have a unique solution to your problem. So you got three rows and you got three columns, and you have all these identity—so these unit vectors—then this is the indication that you got a unique solution to your problem, and that solution is simply this column, as you can see in here. So next time when you have something like this, you don't even need to rewrite the transformation back to the equations; you can directly say that the solution to your system with these three equations and three unknowns is this one—so this column Vector equal with this entry 16, -4, and -6. And that's the end of this problem. So in this way, we have learned how we can quickly solve this type of problems by just following this uh common procedure, and this ended up with a unique solution. So you can have cases when you will not have any solution, and that will be the case when we saw before, two. So in case um there is an equation says that um in the left hand side you have zeros, but then in the right hand side you have a number, so you end up with 0 is equal to two numbers, let's say eight. This this cannot happen, which means that you are saying your system doesn't have any solutions. So let's let me actually write it down. So you can have three possible cases: Case one, when you end up with one of your um equations, and here you got numbers, let's say 3, 4, 5, and here this is zero. So basically you end up saying 0 is equal to 5, which is not something that is possible, so you say the solution, solution is this, so it doesn't exist. It is also possible that you have Case two, when um like in the previous case that we saw before in this example, you don't have um the um you don't have all the uh unknowns expressed in terms of the other ones, which means that you don't have a unique solution to your system. So you can even see that from here; therefore, you have your X2 and X4, so you describe your final solution as a linear combination of some of the uh variables that can take infinite amount of uh values. So X2 can be anything, and X4 can be anything, which means that your Solutions will also be infinite. So this Case two will have infinite amount of solutions, and the Case three is the case that we saw before right now, when we end up with this identity Matrix in here and this corresponding uh amounts in here. So this gives us indication that we got just single solution, so single solution, a solution to your general linear systems.