📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

A Simple Introduction to Copulas

Dirty Quant16:54

Transcription

Today, we're gonna finally understand what copulas are. Hey everyone, Tino here. Welcome to the channel. Thank you for joining me.

So, this is one that obviously bugged me for a long time. You know, I was uh really intrigued by copulas, sort of maybe I said 10, 15 years ago, and I always sort of struggled to understand what they did really. Um, I guess most of the literature is just really, for both, it's just we know lots of Greek letters. Hello, do you want to play a game? Uh, really confusing actually. So, it took a long time for me to really understand, uh, you know, grasp exactly what was going on. So on, so um, once I did, which is actually they're so simple really, um, I thought, okay, let's just uh let's do a video. Let's try and help everyone else out. So, out. So we're gonna start, you know, look at some uh some graphics and essentially go from there.

So, as always, if this is the first time here, my name is Tino. I cover stats, maths, statistics, finance—all good things are that, that—and you know the drill. Like, subscribe. All right, let's fire up a Jupyter notebook. Notebook. Boom.

All right, so um, copulas. So maybe before introducing copulas, just copulas, just have a look at some data and uh really what we what we tend to, you know, do in terms of [Music] [Music] assumptions, right, when we are use something like correlation, we're making some strong assumptions. So let's have a look. So all I'm going to do here is and just I'm going to fix the seed so that that it's all very reproducible. I have a little little uh correlation matrix so with a mean of zero, zero, a row, so correlation coefficient of 0.8, and then I'm actually going to use this uh multivariate normal, np.random.multivariate_normal, generate a thousand data points. Right, I'm going to save that to norm1, num2. I'm also going to transform to uniform, but that will come later on. So I'm just going to run that, and if I run the correlation of that, you can see that they are 0.8. Right, so they are indeed as as I wanted them. And what we can do, just have a quick scatter plot, and this is what the data looks like. So you've got, you know, your typical uh histogram there, um, or which is essentially normal, another one normal, and this is, you know, it's just above at normal. There's not much, there's not much, much really going on here. You've got um I've got this sort of OLS line going through it, and this is really what you expect, right? So really mostly bunched up in the middle, which uh this is where it occurs here. So you've got a big lump there, another big lump there, and that's why you've got that most of the scatter plots happening in the middle.

So, when we're doing, when you're using correlation, we are assuming that this is what data looks like, looks like, right? So, a, it's it's it's linear. There's a linear relationship between them, which means, you know, a line, a straight line is the best thing for it, right? Um, and sometimes that might not be the case. Uh, data isn't always normal, and cover it another time, this misconception that, oh, if I accumulate more data, it's going to become normal, it's like, no, it doesn't actually work like that. Um, that sort of was misused quite a lot. So, so this case, yes, it's a barbaric normal. Correlation is correlation is, you know, linear correlation is the best thing for it, so just a straight line, and that's pretty much it.

So what if the data looks different? And this is pretty much what copulas allow you to do, right, is to really change what uh these edge distributions, so these are called marginals. You'll hear something to do with marginal distribution. So there's really three components, components. One is the marginal, so essentially what goes on the edges. What does that edge distribute? Forget, forget the other one. Let's just look at this top one, right, right, this top one here. Is it normal? Is it not? This is student t, gamma, beta, pick your distribution, right? It can look however you want, and the same for the other one. You don't, they don't have to be the same, and they don't have to be normal, right? So that's the first step, which copulas allow you to decouple, decouple uh from each other. And the third piece of the puzzle, how do they interact together, right? So what is the structure that essentially best models how one um one piece of data interacts with the other one, right? So in this case, it is actually uh so multivariate normal is the best thing for it, so a Gaussian distribution is actually the best thing for it, and I forced it to be like that. But let's see what happens, what happens if that is not the case, right?

So I'm just going to make up some dummy data here. Um, I've got an example, look, example, look, you're on Amazon, your favorite shopping website, and website, and you want to see what's the relationship between how much time you spend on the site, site, and how much money you spend on the site, right? So let's just say you've got here time spent on site. So what I'm actually gonna do here is have this a really nice distribution here. It's actually a gamma distribution. So I use this line up here, all right, gamma. I've got some inputs. It doesn't really matter. I'm just trying to do something that's not your typical, typical bell-shaped sort of Gaussian distribution, and uh yeah, pretty much like this. So you've got most of the essentially most of the people bunched up around here. So time spent is louder than someone like in the five minutes sort of thing, sort of five minutes spent on the website, and there is, you know, one observation up here. There's one person that spent 50 minutes on uh on Amazon. And if they spend more time on Amazon, do they spend more money? Right, that's really what we're trying to want to try and understand, right? Um, so this is this just time spent on website, and website, and you know, it's a nice distribution, something slightly different, right? What about the dollars? How much money you actually spend on the website? And then we've got something like this, you know, really, really completely different to your typical bell distribution. You've got a big spike here and a big spike here. All right, so there's a lot of people that go on Amazon and they don't spend any money at all, just have a browse and they're off, and that's why you've got the big spike here. But at the same time, there's actually a lot of people that spend like, let's just say 100 bucks plus, right? Um, but just go in and just buy something really expensive, and then there's all these sort of uh different observations in the middle. So just a nice sort of U-shaped curve or something like that, uh, and this and this is actually a beta distribution with these parameters here, right? So completely different uh distribution, something slightly different, different. What do they look like together, right? So if I were to use this scatter plot between them, right, so this is essentially what it looks like. You've got that gamma, that beta, and it's yeah, it's really quite, quite interesting because you've got, you know, I've got this sort of line of best fit through it, and is it best? I don't know. It doesn't really account for anything to the right. Um, all these dots here are completely missed out, missed out. And if it were to actually calculate the correlation between them, it comes out sort of, you know, 0.72, and remember I've actually fixed this data, this data at 0.8, so I know that the true correlation between them is 0.8. What is going on? The issue here is I'm trying to use the wrong tool for the job, the job, the the relationship here is not linear, first of all, and secondly, the marginals, what's on the edges, is not normal. So obviously this is gamma, this is beta, uh, and that's why you get those, and this is what, point one off, but you know, this could have looked any, this distributions could have been even, you know, wilder if you wanted, and that number could have really drifted off, right? So your relationship, relationship isn't quite there, and there's also other, other nuances of the data which we'll cover, um, cover, um, copulas capture.

Let's just say for example, like, example, like tail dependence, right? So it might be that two assets don't have a really strong correlation in the middle, so you know, during day-to-day, so financial assets, right? So during the day-to-day, they don't sort of behave in a very correlated way, but when there's maybe a strong market movement, they both correlate. So the tails are actually really co-dependent. They move, they move uh, say they're both down minus five, minus ten percent, they are, that's going to happen at the same time, very likely that that's going to happen at the same time, right? But if you were to look at sort of the daily, sort of normal day-to-day correlation, what's going on, one's going down, it does, there's not really much going on. So copulas allow you to sort of model that dependency as well, right?

So what do you do about it? What if you have something like this? You say, look, I want to use a better tool for it. So this is what I'm talking about, that decoupling, right, of these marginals. So what do I do? Copulas is really uh an exercise in firstly, firstly identifying what the marginals are. In this case, I clearly know what they are because I set them, so um, but there's sort of a bit of um so there's a bit of expertise, I guess, that is is required, that is is required, and there are a few programs out there which try, so you give them some data, and they're trying to work out what they are. I guess you could do a bunch of tests. It's up to you really, but I think plotting it using your eyes is probably the best thing for it. So what you need to do is identify what the marginals are, right? So once if you know what the marginal is, what you can do is use something called a CDF or a cumulative distribution function, right? What does that even mean, right? So what you can actually do, let me just plot it, and that will become very, very clear. CDF allows you to map the input observations from from your data into a space that goes from between zero and one. So uh the lowest value that you you observe will be a zero, and the highest value is one, and this is this little curve that you observe is like sort of like an S, I guess, right? Uh, this curve is that transformation. It's taking an input value from the bottom, so let's say 10. If you look at 10, we go up, and it transforms it, transforms it to just approximately like 0.6, right? So you can see the original and CDF there, and that's what the the that's what this this this function does. It takes an input variable in its original space, so in this case uh time, so there's a minute, transform those minutes into a zero one distribution. You think, well, why am I doing that? The reason for that is that it transforms into a uniform distribution, so distribution, so um I could have had more data points, but approximately this on the right is actually a actually a uniform. I think this, yeah, this plot will show it slightly better. This is essentially a uniform distribution. So it can, it should be completely flat. If I were to sort of assimilate this for 100 million data points, it would be completely, completely flat, right? That's a great place to be, right? I've got my, it means I've identified my marginal, marginal really well. I'll identify what is that distribution really well, and I've got this uniform distribution, which is like e to everything, right?

Let's do the same thing, thing for the dollars spent on the website. In that case, I know it's a beta distribution, so I'm going to use the beta CDF, beta CDF, to again transform that data. So just as before, before, it takes those inputs between, you know, whatever uh whatever uh your dollars are spent on the website and transforms them, so 40 dollars becomes approximately 0.4, right? So um you've got that transformation that maps it from it from dollars to this 0 1 space, and again it should be completely uniform, completely flat, completely flat. Let's plot this one here. Should be a little bit easier to see exactly, right? So so dollar spent insight, completely flat, right? Okay, right, okay.

All right, so what happens if I plot now these two, these two uniform distributions against each other, right? So I'm just going to plot it here, and this is what I get, right? So I've got uniform here, uniform there, and I get this nice sort of, you know, the line, line works out pretty well, and you know, I could, I guess, potentially now that I've got it in uniform, so let's actually have a look. The correlation is that's it, you know, 0.79, I pro, approx, approaching is that 0.8 again. There's going to be a margin of error here because this is again simulated data, which has um by its very nature some sort of random, random element to it, right? But I've gone from something that was distributed completely differently than, you know, you know, my assumptions made which was um, you know, beta and gamma distribution. I've then transformed those, identified what those marginals are, what are those um what are those distributions on their own, transformed into a uniform distribution, and then I I plotted them through this. I haven't actually technically fitted a copula to this, so this so this is like really simple, simple data. Um, I don't really have to, I guess the the concept is there. So it's really an exercise in a identifying what those uniforms are and then, you know, identifying, you know, what the relationship is here. I mean, this is was actually generated uh from multivariate normal, so it's just gonna be just a straight line, right? There's nothing really special, um, but so the last piece which I think is beyond sort of, you know, a nice introduction, keep it nice and sweet, is what different relationships can I have, and have, and let's have a quick look. These are examples of copula, uh, cochlear. Gaussian is essentially what we just had above, right? So that's normal. Gaussian is exactly it, so um more in the middle, less towards the edges, right, edges, right? Um, you can have t, Frank, so you know, interesting ones are something like like a like um Joe, right? Joe has this sort of uh tail dependence at at the right-hand tail, right? I guess you could even have, and there, I think I won't say hundreds, but dozens and dozens and dozens of dozens of copulas, all having slightly different characteristics, different characteristics, and you know what the relationship is in the the in the tails. Um, so this one here, Joe copula has a really strong codependence and the right-hand tail, right? So you see sort of in when when the market, let's say, you know, when the market's going down, they don't have really much of a correlation on the way back, way back up. These two items tend to really sort of of uh behave in a very similar manner, right? So all of these have slightly different ways of uh of behaving and behaving, and again a bit of an art uh you can do various tests, various tests, try and identify which ones which, but that's really it.

So in its basic form, just getting data, transforming into a a common language, that uniform, that uniform data, and actually from that uniform, you can transform it to any other distribution of that that you want, which is the is the which is awful. It doesn't have to be kept in that in that uniform space, right? So when I actually went and generated this data, uh I actually went from, you know, normal to I went from normal to uniform using the CDF, the CDF, right? Uh, but then I actually went from um from that uniform back to that that distribution. I went from say uniform to gamma because uniform is like your jumping spot to anything, anything, right? And I use this and I use this this uh little function here, this this PPF, it's like an inverse CDF, right? Uh inverse CDS, inverse CDF function, which allows you to flip from uniform, uniform to the distribution and distribution to uniform is just normal EDF, right?

Um, I hope you enjoyed the video, video. Um, there's so much more that we can cover with, cover with with copulas, but in terms of a quick introduction, I hope this uh this um this is clear. It's a real deep and complex part of statistics and maths, and a lot of people don't tend to cover it because the literature is the the formulas are brutal. There's a lot of proofs and etc. But in this high level, um, I think this example, example is pretty clear. Catch you in the next one, one. [Music] [Music] Cheers