Transcription
In the last videos, I talked about the derivatives of simple functions. The goal was to give you a clear picture or intuition to understand where these formulas come from. However, most functions you encounter when modeling the world involve mixing, combining, or tweaking these simple functions in some way. So, our next step is to understand how to take derivatives of more complicated combinations.
Again, I don't want you to memorize these; I want you to have a clear picture in mind of where each one comes from. This really boils down to three basic ways to combine functions: you can add them together, you can multiply them, and you can compose them, which means putting one function inside another. Sure, you could say subtracting them, but that's just multiplying the second by negative one and adding them together. Likewise, dividing functions doesn't really add anything, because that's the same as plugging one function inside another and then multiplying the two together.
So, most functions you come across involve layering these three different types of combinations, though there's no limit to how complex things can become. But as long as you know how derivatives interact with just those three combination types, you'll always be able to take it step by step and peel through the layers for any kind of complex expression.
The question is, if you know the derivative of two functions, what is the derivative of their sum, their product, and the composition of the two? The sum rule is the easiest, if somewhat tongue-twisting to say out loud. The derivative of a sum of two functions is the sum of their derivatives.
It's worth warming up with this example by really thinking through what it means to take a derivative of a sum of two functions, since the derivative patterns for products and function composition won't be as straightforward and will require deeper thinking.
For example, let's consider the function f of x equals sine of x plus x squared. This function adds together the values of sine of x and x squared for every input. For instance, at x equals 0.5, the height of the sine graph is represented by this vertical bar, and the height of the x squared parabola is represented by this slightly smaller vertical bar. Their sum is the length you get by stacking them together.
For the derivative, you want to ask what happens as you nudge that input slightly, maybe increasing it to 0.5 plus dx. The difference in the value of f between those two points is what we call df. When you picture it like this, you'll agree that the total change in height is whatever the change to the sine graph is, which we might call d sine of x, plus whatever the change to x squared is, dx squared.
We know that the derivative of sine is cosine, and remember what that means. It means that this little change, d sine of x, is about cosine of x times dx. It's proportional to the size of our initial nudge dx, and the proportionality constant equals cosine of whatever input we started at. Likewise, because the derivative of x squared is 2x, the change in the height of the x squared graph is 2x times whatever dx was.
So, rearranging df divided by dx, the ratio of the tiny change to the sum function to the tiny change in x that caused it, is indeed cosine of x plus 2x, the sum of the derivatives of its parts.
But as I said, things are a bit different for products, and let's think through why in terms of tiny nudges again. In this case, I don't think graphs are our best bet for visualizing things. Commonly in math, if you're dealing with a product of two things, it helps to understand it as some kind of area.
In this case, you might visualize a box where the side lengths are sine of x and x squared. But what would that mean? Since these are functions, you might think of those sides as adjustable, depending on the value of x, which you can freely adjust up and down.
To get a feel for this, focus on the top side, which changes as the function sine of x. As you change this value of x up from 0, it increases to a length of 1 as sine of x moves up towards its peak, and after that, it starts to decrease as sine of x comes down from 1. Similarly, that height is always changing as x squared.
So, f of x, defined as the product of these two functions, is the area of this box. For the derivative, let's think about how a tiny change to x by dx influences that area. What is the resulting change in area df? The nudge dx caused that width to change by some small d sine of x, and it caused that height to change by some dx squared.
This gives us three little snippets of new area: a thin rectangle on the bottom whose area is its width, sine of x, times its thin height, dx squared; a thin rectangle on the right, whose area is its height, x squared, times its thin width, d sine of x; and a little bit in the corner, which we can ignore. Its area is ultimately proportional to dx squared, and as we've seen before, that becomes negligible as dx goes to zero.
This whole setup is very similar to what I showed in the last video with the x squared diagram. Just like then, keep in mind that I'm using somewhat larger changes here to illustrate things, just so we can actually see them. But in principle, dx is something very small, and that means that dx squared and d sine of x are also very small.
Applying what we know about the derivative of sine and of x squared, that tiny change, dx squared, is going to be about 2x times dx. And that tiny change, d sine of x, will be about cosine of x times dx. As usual, we divide out by that dx to see that the ratio we want, df divided by dx, is sine of x times the derivative of x squared, plus x squared times the derivative of sine.
Nothing we've done here is specific to sine or x squared. This same reasoning would work for any two functions, g and h. Sometimes people like to remember this pattern with a mnemonic that you can kind of sing in your head: left d right, right d left.
In this example, where we have sine of x times x squared, left d right means you take that left function, sine of x, times the derivative of the right, in this case, 2x. Then you add on right d left, that right function, x squared, times the derivative of the left one, cosine of x.
Now, out of context, presented as a rule to remember, this might feel pretty strange, don't you think? But when you actually think of this adjustable box, you can see what each of those terms represents. Left d right is the area of that little bottom rectangle, and right d left is the area of that rectangle on the side.
By the way, I should mention that if you multiply by a constant, say 2 times sine of x, things end up a lot simpler. The derivative is just the constant multiplied by the derivative of the function, in this case, 2 times cosine of x. I'll leave it to you to pause and ponder and verify that makes sense.
Aside from addition and multiplication, the other common way to combine functions, and believe me, this one comes up all the time, is to shove one inside the other: function composition. For example, we might take the function x squared and shove it inside sine of x to get this new function, sine of x squared.
What do you think the derivative of that new function is? To think this through, I'll choose yet another way to visualize things, just to emphasize that in creative math, we have lots of options. I'll set up three different number lines: the top one will hold the value of x, the second will hold x squared, and the third will hold the value of sine of x squared.
That is, the function x squared gets you from line 1 to line 2, and the function sine gets you from line 2 to line 3. As I shift this value of x, maybe moving it up to the value 3, the second value stays pegged to whatever x squared is, in this case moving up to 9. The bottom value, being sine of x squared, will go to whatever sine of 9 happens to be.
For the derivative, let's again start by nudging that x value by some little dx. I always think it's helpful to think of x as starting at some concrete number, maybe 1.5 in this case. The resulting nudge to that second value, the change in x squared caused by such a dx, is dx squared.
We could expand this like we have before, as 2x times dx, which for our specific input would be 2 times 1.5 times dx, but it helps to keep things written as dx squared, at least for now. In fact, I'm going to go one step further and give a new name to this x squared, maybe h, so instead of writing dx squared for this nudge, we write dh.
This makes it easier to think about that third value, which is now pegged at sine of h. Its change is d sine of h, the tiny change caused by the nudge dh. By the way, the fact that it's moving to the left while the dh bump is going to the right just means that this change, d sine of h, is going to be some kind of negative number.
Once again, we can use our knowledge of the derivative of sine. This d sine of h is going to be about cosine of h times dh. That's what it means for the derivative of sine to be cosine. Unfolding things, we can replace that h with x squared again, so we know that the bottom nudge will be a size of cosine of x squared times dx squared.
Let's unfold things even further. That intermediate nudge dx squared is going to be about 2x times dx. It's always a good habit to remind yourself of what an expression like this actually means. In this case, where we started at x equals 1.5 up top, this whole expression tells us that the size of the nudge on that third line is going to be about cosine of 1.5 squared times 2 times 1.5 times whatever the size of dx was.
It's proportional to the size of dx, and this derivative gives us that proportionality constant. Notice what we came out with here. We have the derivative of the outside function, still taking in the unaltered inside function, and then multiplying it by the derivative of that inside function.
Again, there's nothing special about sine of x or x squared. If you have any two functions, g of x and h of x, the derivative of their composition, g of h of x, is going to be the derivative of g evaluated on h, multiplied by the derivative of h. This pattern is what we usually call the chain rule.
Notice for the derivative of g, I'm writing it as dg dh instead of dg dx. On the symbolic level, this is a reminder that the thing you plug into that derivative is still going to be that intermediary function h. But more than that, it's an important reflection of what this derivative of the outer function actually represents.
Remember, in our three-line setup, when we took the derivative of the sine on that bottom line, we expanded the size of that nudge, d sine, as cosine of h times dh. This was because we didn't immediately know how the size of that bottom nudge depended on x. That's kind of the whole thing we were trying to figure out.
But we could take the derivative with respect to that intermediate variable, h. That is, figure out how to express the size of that nudge on the third line as some multiple of dh, the size of the nudge on the second line. It was only after that that we unfolded further by figuring out what dh was.
In this chain rule expression, we're saying to look at the ratio between a tiny change in g, the final output, to a tiny change in h that caused it, h being the value we plug into g. Then multiply that by the tiny change in h, divided by the tiny change in x that caused it.
Notice that those dh's cancel out, giving us a ratio between the change in that final output and the change to the input that, through a certain chain of events, brought it about. That cancellation of dh is not just a notational trick; it genuinely reflects what's going on with the tiny nudges that underpin everything we do with derivatives.
So, those are the three basic tools to have in your toolkit for handling derivatives of functions that combine many smaller things. You've got the sum rule, the product rule, and the chain rule.
I'll be honest with you: there is a big difference between knowing what the chain rule is and what the product rule is, and actually being fluent in applying them in even the most complex situations. Watching videos about the mechanics of calculus will never substitute for practicing those mechanics yourself and building up the skills to do these computations.
I really wish I could do that for you, but I'm afraid the ball is in your court, my friend, to seek out the practice. What I can offer, and what I hope I have offered, is to show you where these rules actually come from. To show that they're not just something to be memorized and repeated, but natural patterns—things that you too could have discovered just by patiently thinking through what a derivative actually means.