📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Lec 12: Higher order differentiations

NPTEL IIT Guwahati48:50

Transcription

Welcome to another lecture of this course called Mathematics for Economics, part one. So, in the previous series of lectures, I think there were three lectures, we covered the basics of Differentiation. So, we defined what is the derivative of a function and what is the Newton quotient. We had discussed a bit about the limits, because the idea of limits is important to define differentiation and derivative. We also talked about rules of limits and rules of differentiation. So, what happens if you have the addition of two functions? If you take the derivative of that, then what do you get? Or if you have the difference of two functions, and if you take the derivative of that, then what do you get? Is the product of two functions, and that we saw is called the product rule. We also talked about what is known as the quotient rule. If you have the quotient of two functions, and you want to take the derivative, then what is the result? Finally, we looked at partial differentiation. Because it may happen that a function has more than one variable as the independent variables. And if you want to find out how the function changes with respect to changing each of these independent variables, then we get what is known as Partial Differentiation. So, those things have been covered.

So, today we shall start with a related topic, and this related topic is called Higher Order Differentiation and Linear Approximation. So, this is the first slide: that the derivative of a function is called the first derivative or f dashed of x; this is the same as f dashed of x = d/dx of f(x). And if you have y = f(x), then you have d/dx of f(x) = dy/dx. So, this is called the first derivative or f dashed of x. This function f dashed x can be further differentiated to get the second or higher order derivatives. So, if you have f dashed x, then f dashed x itself is a function of x; f(x) is a function of x itself, but if you take the derivative with respect to x, you get f dashed of x. Now, f dashed of x itself is a function of x. So, it can be differentiated further with respect to x. And if we do differentiation of f dashed of x once, then we get the second derivative. The second derivative is denoted by f double dashed of x. So, you have 2 dashes here, so, we are calling it as the second derivative or double derivative of f(x). There are other ways to denote the same thing: f double dashed x = d²f(x)/dx² or it is the same thing as y double dashed and which is the same thing as d²y/dx². If we assume y = f(x). So, there are different ways to denote the same thing.

Now we take one example. Suppose this function is given: f(x) = 3x² - 4x + 6. So, this is equal to f(x), so we have to find the first and the second derivative of this function. So, f(x) is given: 3x² - 4x + 6, so, we take the first derivative; we use the power rule, which is: if we have x to the power n, and if we want to differentiate that with respect to x, then we get n multiplied by x to the n - 1. So, 3 multiplied by 2 multiplied by x. So that becomes 6x and -4x. So, if you differentiate x with respect to x, you get 1. So, finally, you are getting 6x - 4, because the derivative of 6, which is a constant, is 0. So, this is the first derivative: df(x)/dx = 6x - 4. Now, we want to find out what is the second derivative? So, for that, we have to differentiate this once again with respect to x, and if we do that, we shall get d²f(x)/dx² = 6. The reason being that -4 will drop out because the derivative of a constant is 0. And if you want to differentiate x with respect to x, you get 1 and 6 is a constant. So you have six multiplied by 1; therefore, you have only 6.

If the second derivative is differentiated once more, then what happens? We get the third derivative. So, this way, it goes on in an iterative manner. So, in the above example, if you have f(x) = 3x² - 4x + 6 and from this we had obtained the second derivative to be 6. So, if I want to find the third derivative, I have to differentiate 6 with respect to x and that is equal to 0 because the derivative of a constant is 0. So, the third derivative in this case is equal to 0. Now, this third derivative is also written as f⁽³⁾(x) instead of these 3 dashes, I just write 3 within brackets. So f⁽³⁾(x) or it could be written as d³f(x)/dx³. So, this is the third derivative. And likewise, we can go on the order of the differentiation and we get the 4th derivative f⁽⁴⁾(x), the 4th derivative or d⁴f(x)/dx⁴, and this way it can go on and we can write it for any arbitrary number n, n is any positive integer and so it will be fⁿ(x) or dⁿf(x)/dxⁿ; this is called the nth derivative of the function f(x). This procedure of obtaining higher order derivatives can be performed for partial derivatives as well.

So, this was so, in the case where f(x) was a function of just one variable x, y = f(x), but if your function has more than one independent variable, then we know we can take the partial derivatives and for partial derivatives also one can find out the higher order derivatives. So, let us take one example. General case: suppose y = f(x₁, x₂, ..., xₙ) is a function of n independent variables and these variables are x₁, x₂ … xₙ. Now, we know one can obtain n partial derivatives of first order from each of the n independent variables. So, this is written as ∂y/∂xᵢ, where i can take any of these values. So, ∂y/∂x₁ could be the partial derivative with respect to x₁, then you have ∂y/∂x₂, ∂y/∂x₃ etc., etc., ∂y/∂xₙ. So, n partial derivatives could be obtained and all these derivatives are of the first order. The same thing is written as f’ᵢ(x₁, x₂, …, xₙ) and the same thing can be written as ∂f/∂xᵢ. So, n partial derivatives of the first order are possible. For each of the n first order partial derivatives, one can obtain n partial derivatives of the second order. So, this is how it is written. So, you are starting from this: ∂f/∂xᵢ. So, you are differentiated partially this function with respect to xᵢ, and this thing ∂f/∂xᵢ, it could be differentiated with respect to each of the independent variables and let us assume that xⱼ is that arbitrary independent variable. So, the expression that you are going to get is, is going to look like this: ∂/∂xⱼ(∂f/∂xᵢ). So, two partial derivatives have been taken: first with respect to xᵢ and then with respect to xⱼ. Mind you, i could equal to j also, but it is not necessary that i = j; i and j are any 2 arbitrary numbers less than or equal to n. So, this expression is written as also as y’’ᵢⱼ or f’’ᵢⱼ(x), where, if you notice this x that I have written here is boldface x, which basically denotes this vector. So, this vector. So, this is the second derivative.

Now, how many secondary derivatives are possible from this function? In all there will be n multiplied by n, that is n² second order partial derivative terms and this is easy to see: from each of the n independent variables you get one partial derivative of the first order and from each of the first order partial derivatives, you can further differentiate partially with respect to n independent variables and so, from each of the n you get n more second order partial derivatives. So, in total n multiplied by n = n² such second order partial derivative terms are obtainable. And this n² terms are arranged in an n by n matrix and this matrix has a specific name: it is called the Hessian Matrix and this is how this Hessian matrix is written: you can see this matrix has n rows and n columns. So, each of these rows: suppose you take the first row, the first number in this subscript is 1. So, first the f function has been partially differentiated with respect to 1 and so, you will get just one term. So, f₁₁(x) then this f₁₁(x) is differentiated partially second time with respect to each of these n variables. So, you are getting this whole range of numbers: n numbers and this is performed for each of these variables. And you are going to get n rows n columns. So, this is an n by n matrix and within the brackets these are the x’s and these x’s are as we have seen these are vectors. So, this Hessian matrix is evaluated at this particular vector x, where x = (x₁, x₂, …, xₙ). So, when we are dealing with a function with more than one variable and we are doing this, this operation of partial derivatives then we talk about second order partial derivatives then this idea of a Hessian matrix becomes very important because the second order partial derivatives are represented through this Hessian matrix.

Now, we talk about something else which is called the Chain rule of derivatives. Suppose y is a function of the variable u. So, y is a function of u, but u itself is a function of x. So, y is called a composite function of x, because there are two functions involved here: one is the f function which is a function of u and the other function is the g function which is a function of x. So, therefore, is called a composite function; it involves more than one function of x. In this case the change in x sets off a chain reaction first to u and via u to y. This is captured by the chain rule. So, this is easy to see: suppose x is changing, then this is going to affect u, u is going to change and therefore, if u is going to change that will affect y. So, a kind of chain reaction is a set of: if x changes and the effect of that chain reaction ultimately reaches y. So, this chain reaction is captured by this chain rule which is written here: dy/dx = (dy/du) * (du/dx). For the chain rule to hold both y and u have to be differentiable functions: that is very obvious, because you have to be able to differentiate g with respect to x and f with respect to u, only then you can talk about these terms on the right hand side of the chain rule. In other words, what does it mean? It means that the rate of change of y with respect to x is the product of the rate of change of y, this should be change. Rate of change of y with respect to u multiplied by the rate of change of u with respect to x: that is what we have written: rate of change of y with respect to x = rate of change of y with respect to u multiplied by the rate of change of u with respect to x.

So, let us take one example to focus our ideas. So, suppose x is given as a function of small t. So, the specific form is this except kind of complicated form: 10 * (1 + √(t² + 1))¹⁵. So, it is a function of t but it is a little bit complicated function. So, we have to find out what is x’(t)? So, the derivative of x with respect to t that we have to find out, this dx/dt. Now, we apply the chain rule, but before that let us define this function u. So, u: suppose this expression inside the brackets is 1 + √(t² + 1). So, if I take u = that term, then x(t) becomes 10 * u¹⁵ by the given function of x and we apply the chain rule that x’(t) that is dx/dt = (dx/du) * (du/dt). So, what is dx/du? So that we can obtain from here, it will be 10 * 15 * u¹⁴, I am using the power rule that if you take the derivative of xⁿ, you get n * xⁿ⁻¹ and there is another term which is du/dt. So, how does the u change with respect to change in time? So here t can be assumed to be time. So, that is what we have written here. We have just substituted the value of u from here, u was assumed to be this. So, I am using that expression of u which is 1 + √(t² + 1) that we have to differentiate with respect to t. And this is preceded by 10 * 15 which is 150 * u¹⁴. And now actually we have to use another function, another chain rule will apply to simplify this expression, we assume that v = t² + 1. So, within the root over we have now v instead of t² + 1 and we once again apply the chain rule. So, d/dt of this thing will be d/dv of this thing multiplied by dv/dt and d/dt of this expression will be d/dv of √v because 1 will cancel, 1 is a constant multiplied by dv/dt and again I use the chain rule which is you have v½ here, root over is nothing but ½. So, ½ comes first and then v⁻½, ½ - 1 = -½ and which can be written as 1/√v multiplied by dv/dt. Now, we have to find out what is dv/dt? dv/dt is not known to us. So, from here, I can find that here if I differentiate v with respect to t, again I use the power rule and I will get 2t, and I substitute that back in this expression, and I get this x’(t) = notice what I am doing here: 150 * ½ = 75 and then we have u¹⁴. So, I have just substituted back the expression for u here. So, u¹⁴ multiplied by 1/√v. So, this is √v and multiplied by dv/dt which is 2t. And this can be written in this form because 2 * 75 = 150 and I have taken t in front and the rest of the terms are the same. So, you have this to be our answer. Now, here what we have done is actually we have extended the chain rule one step more, we have actually used this kind of formula that dy/dx = (dy/du)(du/dv)(dv/dx). First we have used the u function with proper definition and then we have used the v function here with a proper definition and there we caught the expression for dx/dt.

Alternatively, the chain rule can be written as the following: suppose y is a function of x and you have x is itself a variable but it affects u. So, y is a function of x can be written as f(u(x)) because u(x) affects y, but x itself affects u. So, I write this as this: a function of a function where y is a composite function with f as the exterior and u as the kernel, kernel means something which is inside and f is that function which is outside. So, you have basically two functions fused together. From this we write that y’(x₀) = f’(u(x₀)) * u’(x₀). So, you have basically two derivatives like we have in the chain rule. First f is differentiated with respect to u, but u is evaluated at x₀ and multiply that with the derivative of u with respect to x where x is evaluated at x₀.

Now, we come to another kind of rule which is called Implicit Differentiation. So, we start with an example: suppose this is given: x * √y = y⁴. Here it is not explicitly mentioned y is a function of x. So x and y are merged together. So we have to find out what is y’, that is dy/dx. So what we do is that we take the derivative of both sides with respect to x, and use the product rule and chain rule. So if I take the derivative of both sides, and I use the product rule first. So, the first function which is x, then multiplied by d/dx of the second function, which is √y, plus the second function √y multiplied by the derivative of the first function, which is dx/dx. And on the right hand side, you have the constant. And if you take the derivative of the constant it becomes 0. And then what do we do, we basically now use the chain rule. So there we have used the chain rule, first we differentiate √y with respect to y and multiply that with dy/dx. And the other terms are kept intact. Now, if I take the derivative of √y with respect to y, I get this expression of 1/√y, multiplied by dy/dx. And the second term will just be √y, because d/dx of x = 1. So, the √y multiplied by 1 is √y, and on the right hand side you have 0. And then I take dy/dx to one side, which is what I need to find, which is y’. And on the hand side, if I simplify this, it will just be -2y/x. So this is what I was supposed to find: y’. So, this was the first derivative dy/dx, but from this first derivative, actually, I can find the second derivative. But notice how we have proceeded so far, although the function is not explicitly mentioned, y is not mentioned as equal to f(x), it is implicit, the y and x are marched together. So, implicitly, y is a function of x. And we have just taken the derivative of both sides with respect to x and proceeded. And we have found that dy/dx. Now, dy/dx is this. So I can find out from here the second derivative of y with respect to x by differentiating this expression -2y/x with respect to x. And if I do that, I use the quotient rule now, so - of quotient, which is a ratio on the denominator I have x², the square of the denominator on the numerator, I have the denominator multiplied by the derivative of the numerator. So, 2y that I differentiate with respect to x, I get 2y’ - the numerator multiplied by the derivative of the denominator, which will give me just 1. And what is y’? So, y’ is substituted from here. So I include that here. And if I simplify this a bit, I will get this expression, 6y/x².

So, before we conclude this lecture, here are some examples from economics, where I use implicit differentiation and also chain rule. So, here is the first example. So, you have C, which is a function of Y. Remember how we named this function before, this is called a Consumption function. So, this is in the context of macroeconomics, you have the total consumption of the economy, which is called the consumption function and it is a function of the income of the economy. So, I have taken a very simple consumption function: linear: 100 + 0.7 * y. And we know this identity that y = C + I. Y is output or income, which is equal to consumption plus investment. So, we have to answer three questions. First, find Y as a function of I. Find dY/dI and find dY/dI for a general function C = f(Y). Let us go 1 by 1. So we substitute this particular function which is given to us, the consumption function into this macroeconomic identity, y = C + I. And if I do that, I get this: y = 100 + 0.7 * y + I. So, then I have to find out Y as a function of I. So what do I do? I take Y to one side, therefore, I will get on the left hand side, 1 - 0.7 = 0.3, 0.3y = on the right hand side you have 100 + I. So, therefore, I take, I divide both sides by 0.3, and I will get this particular expression for Y: y = 1000/3 + (100/3) * I. Now, from the above, how do I find out that dY/dI? I can use this explicit function, this is not even an implicit function, this is an explicit function of Y. And I differentiate this explicit function of Y with respect to I, I will get 10/3. That is the answer. And the third part is: suppose the form of C is not given: C is just given to be a function of Y: f(Y). And I have to find out dY/dI. Now, how do I do that, I follow the same method as before. I substitute C = f(Y) in the macroeconomic identity. So, I will get this and I take the Y terms to the left hand side, and I get y - f(Y) = I. Now, comes the role of the implicit differentiation, I differentiate both sides with respect to I, because I have to find out dY/dI. And if I do that, I get this term: dY/dI - d/dI of f(Y) = 1, because dI/dI = 1. And then I use the chain rule: dY/dI - d/dY of f(Y) multiplied by dY/dI = 1. So, I have broken this down into two parts. And then what is d/dY of f(Y)? Let us suppose that f’(Y) multiplied by dY/dI = 1. So, now I can take a dY/dI common, And I get dY/dI multiplied by 1 - f’(Y) = 1. And so I divide both sides by 1 - f’(Y) and I get this expression: dY/dI = 1/(1 - f’(Y)). Now, this is a very important expression in macroeconomics, this expression is often called the Investment Multipliers. So, what it tells us is the rate of change of income, which is the GDP of the country if investment expenditure changes by 1 unit so, it shows us the impact of investment on the national income. That is why it is called Investment Multiplier. And these you can see it is a function of f’(Y)? Because f’(Y) is coming in the denominator as a negative term? What is f’(Y)? f’(Y) is the rate of change of consumption with respect to income and that we have seen is called the Marginal Propensity to Consume, MPC: we have talked about this before in another context. So, MPC is something which appears in the investment multipliers and you can verify that as MPC rises, then the value of the investment multiplier also rises. So, what it means intuitively is that if people spend a greater portion of their additional income in consumption, that basically increases the impact that investment increment might have on the national income alright. And we have seen before that this MPC should take a value less than 1 and you can now see why that is true: if MPC = 1 then this term becomes undefined: it becomes 1/0, so MPC to make this term meaningful should take a value less than 1.

I now talk about another example, where implicit differentiation and chain rule will be used: suppose the demand and supply functions in the market are given by D. So, this is the demand function: quantity demanded is a function of price and S: quantity supplied is a function of price: here a, b, α, β these are all positive constants, these are what we have talked about: a is they are called parameters, but what is small t? Small t is appearing here in the demand function; it is not appearing…

In the supply function, mind you, here t is the tax per unit of the good imposed on the consumers of the good. So, if you want to buy 1 unit of the good, then you have to pay this small t, and you are a buyer, so you have to pay the money, not the seller. So, apparently, the government is collecting this tax from the consumers by increasing the price, or not the price is given, but the price that is paid by the consumer that gets increased by this small t. Mind you, the money that the seller will get does not increase; it remains at P. So, the gap between P plus t and P goes to the government as tax.

So, from the above, we can talk about what is known as the market equilibrium condition, and this is just demand is equal to supply: a - b(P + t) = α + βP. Now, we have to again answer three questions: find dP/dT from the above, that is, how much the price changes, how much does the price change with respect to t? So, that we have to find from the above equation. Secondly, we have to solve for P explicitly and find dP/dT. Third, if capital T is the tax revenue, find its expression and find the t which maximizes the tax revenue, alright. So, let us see how it goes.

Now, this is the equilibrium condition, and differentiating both sides with respect to t. So, here we are using basically the implicit differentiation, and if I do that, so, on the left-hand side, smaller is constant and that drops out, -b(dP/dt) - b(dP/dt), which is just equal to 1, and on the right-hand side, you have α differentiated with respect to t; it will become 0 and β(dP/dt). And I take all the dP/dt terms to the left-hand side, and the coefficient of that will be β + b, and on the right-hand side, I take the constant terms, and that will become -b. So, dP/dt, therefore, will be -b/(b + β). I have divided both sides by b + β. So, this is the expression for dP/dt, and this is negative because β and b are positive. That basically means that as the tax rate rises, the equilibrium price falls. The rate of fall is given by -b/(b + β).

You can actually imagine this in terms of a diagram that you have this P, and suppose this is Q, and you have this supply function, this demand function; the demand function is shifting downwards as the t is rising, and therefore, the equilibrium price is declining. So, that is it.

Second part: in the second part, we need to find out explicitly the form of P. So, here implicitly I differentiated and I found dP/dt. But suppose I find out explicitly what is the form of P, the equilibrium price, and then I find out dP/dt. So, from the equilibrium condition, I use the equilibrium condition to find out first the equilibrium price. And if I simplify this, I am just skipping some steps, I get P = (a - α)/(b + β) - (b/(b + β))t. So, this is actually a linear function of t; first, you have a constant term, and then you have some coefficient multiplied by t. From the above, by differentiating both sides with respect to t, we get simply this: -b/(b + β). This is what we have got earlier also by implicit differentiation. Here we did not use implicit differentiation; we explicitly found out the expression for P, and then we got the expression for dP/dt.

Finally, what is the last question? The last question is: if capital T is the tax revenue, find its expression and find the small t which maximizes the tax revenue. So, here we are looking at the situation from the point of view of the government; the government is getting some tax revenue; the total amount of tax revenue that it is getting that is denoted by capital T. Now, suppose the problem for the government is it wants to maximize this tax revenue, and it wants to earn the maximum amount of money that it can get from the market by imposing this tax on the consumers. Now, how does the government then fix the small t? Remember, the government has control over small t; that is the only thing that the government can do. So, the natural question arises: if the government wants to maximize the capital T, which is the tax revenue, then at what value of small t is the capital T getting maximized? So, the government would like to impose that rate of tax, and that will generate the maximum tax revenue. So, therefore, we first find out what is capital T, the expression for capital T, and then try to answer the question: at what small t is it going to get the maximum value?

Now, what is the tax revenue after all? Tax revenue is, after all, the equilibrium quantity of goods that is being purchased and sold multiplied by the tax rate, because the tax rate is imposed on each unit of the good. So, the total amount of units that is being sold and purchased multiplied by small t will give you the total tax revenue. Now, what is the equilibrium quantity? So, it is given by any of the expressions of either demand or supply, but these demands and supplies are evaluated at the equilibrium values. So, equilibrium demand and equilibrium supply. So, I just take the supply expression, which is α + βP, where P is the equilibrium price. Why is P the equilibrium price? Because I want to find out what is the equilibrium quantity that is being supplied and demanded, and that can be found out only at the equilibrium price. So, the equilibrium price is already found; it is this that I have to use here; I have to substitute that here. And mind you, I have just broken down this term. So, it will be α + β[(a - α)/(b + β) - (bt/(b + β))] multiplied by small t. So, our capital T, which is the tax revenue to the government, is this expression; I have to simplify this expression a bit. And basically, I get this by removing the brackets. And therefore, I get this expression: t(α + β[(a - α)/(b + β) - (bt/(b + β))]). Now, the interesting thing to note is that this function, it is a function of small t, so capital T is a function of small t, but it is a function where the form is such that it is t multiplied by another function of small t. So, t is multiplied by something which is itself a function of small t, so it is a quadratic function. So, that is there.

Now, if you look at this function, then this function will attain the value of 0 at two points: at t = 0 and at t = this expression. So, it is like this. So, on the x-axis, you have the small t, which is the tax rate, and on the y-axis, you have the total tax revenue. And if you visualize this, it is like this. So, it is attaining 0 at a point of origin and at this point, which is (αb + aβ)/(bβ), and so it is a basically parabolic shape. So, therefore, the maximum is obtained at half of this. So, it is (αb + aβ)/(2bβ). So, that is why I have written that the maximum will be reached at this middle point, and the middle point is this: (αb + aβ)/(2bβ). So, therefore, we have solved the problem for the government. If the government imposes this small t, the rate of tax, then its total revenue will be total tax revenue will be maximized. So, that concludes the lecture.

So, just to make a concluding remark, in this lecture, we have talked about the higher-order differentials, higher-order derivatives, and then we have talked about the fact that for partial derivatives also, we can talk about the higher-order derivatives, and there is this idea of a Hessian matrix related to that, and then we talked about the chain rule and implicit differentiation, and then we talked about applications of these in practical economic problems. And so, let us conclude it here. Thank you.