📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Why Trading Metrics are Misleading (Unless This is True)

Roman Paolucci35:07

Transcription

[Music] Roman, I have a trading strategy that generates 30% expected returns. Or maybe I'll get something like, I have a strategy that prints money, $3,000 a month. Or or another one, my strategy has a a 95% win rate. Now, all of those performance measures are are largely nonsense, and we'll talk about why that is the case, but let me even take this a step further.

People will come to me and they'll say, "Hey, you know, I have a trading strategy and it has a sharp ratio of 2.7, a sortino of of 3.2, a max draw down of of 3%." And still I say, "Show me the money." That is not at all what I care about when it comes to these performance measures. They are necessary but not sufficient. There's only one thing that I care about when it comes to trading strategies and the relative performance of those strategies. And we're going to talk about what that is in this video.

Before we can discuss this idea, we have to understand what it is. We're dealing with these different trading and investing performance measures. I've separated them into two categories. The first is this idea of useless trading and investing performance metrics. things that your YouTube, Instagram, Tik Tok traders are going to site, things that don't actually have a lot of value, things that will pique your interest but don't actually offer any real insight into a trading strategy's performance.

Now, the second category are metrics that I would consider less useless. These are things that by themselves are not sufficient to suggest that we should create a strategy. But in addition to this idea of stability could suggest that we have a viable trading strategy. A good sharp ratio doesn't mean that you're going to want to trade that strategy. However, a new Quanild video sure does mean that you will have access to a Jupyter notebook. This Jupyter notebook will be linked in the description below. I will also post it to the Quan Guild library on GitHub where you can find all of my Jupyter notebooks and associated YouTube videos along with the source code for all of my Quant builds.

At the top of this Jupyter notebook, you'll notice some related Quanilled videos on ideas applying statistics to trading. The idea that the expectation or some sort of nonlinear conditional expectation is the best we can do in the face of randomness or something that we are constructing as a random variable that may just have too many deterministic factors to account for. If you're unfamiliar with any of these ideas, I highly encourage you check out these videos before tackling something like this one.

Moreover, these videos certainly take a lot of effort to create. So, if you would like to help support the channel, please like, comment, subscribe. It helps me out tremendously. It is greatly appreciated. And if you'd like to master your quantitative skills, check out quant.com. Maybe you come from a background in business, economics, computer science, and aren't too sure where to start on your quant journey. Quantitative finance is such an interdisciplinary field, combining ideas from math, statistics, probability, finance, computer science, it's hard to figure out where to get started. Well, Quant Guild is for you. We have over 90 quant lessons in math, probability, finance. An adaptive practice engine that scales with your skill level so you can progress to more and more difficult topics with gamified rank-based progress. Interview questions with fully worked solutions, creating games coming soon, courses in math, statistics, finance, coding, all from the ground up from A to Z. And everything's included with Quank Guild membership. So, if you'd like to help support the channel and master your quantitative skills, check out quangild.com.

Without further ado, let's discuss this idea of evaluating investing and trading strategies. In either case, we're going to look at historic data. We're going to either construct some sort of portfolio to evaluate an investing strategy or we're going to run a back test on some sort of trading strategy, come up with the returns, come up with all these different performance ratios, so on and so forth. But we're operating all on historic data, right? I thought historic data, past performance wasn't indicative of future performance. Why are we looking at all of this data in a backward-looking setting if we're trying to assess whether there's anything that is tradable in a forward-looking setting? This is a big piece of the puzzle that we need to find in order to get a complete picture of these metrics and their place in evaluating these strategies.

In efforts to find this piece of the puzzle, let's start with what I would consider useless performance measures. These are expected returns, win rate, cash generated, things like that. Now, we'll start with this idea of expected returns. Expected returns or the average return is easy to compute in a backward-looking sense, but on their own are not a particularly useful measure. We can take any trading strategy, any investing strategy. We can construct a portfolio in an investing sense. We can run a back test in a trading sense. We can take a look at the average return for, you know, daily, weekly, monthly intervals. But this says nothing of the expected return in a forward-looking sense. This is telling us what has been, not what will be. And it's not even telling us what has been in a particularly useful way. It's not even attaching any measure of risk to this relative performance. So if you told me that, you know, you had an expected return of 30%. And another strategy had an expected return of 30%. Well, how do I distinguish between those two strategies? I'm going to assess how much risk you had to take to achieve either return. If we just look at expected returns, we have no notion of risk. I can't possibly evaluate your strategy in any case given just expected return. So it's difficult because it doesn't give us any sort of relative measure to risk for how much risk you had to take for that return. And it doesn't offer any explanability in a forward-looking sense. It doesn't tell us if this is behavior that we're expecting in the future.

A simple argument for why past return isn't indicative of of future performance can be made in a bunch of different ways using say Nvidia. Just this year we're going to be looking at the price path of Nvidia and the expected returns. Now I could look back what one year, two years, five years, 10 years. Is that really appropriate to construct a measure of expected performance? 10 years ago, Nvidia was an entirely different company. This measure is not valid in a forward-looking sense, but it can be even more incorrect if you go too far back. Moreover, if we actually look at this chart, we can see that during the craziness of the tariffs at the start of this year, if we looked at the year-to-date average expected return, we saw something like minus.4% daily. But after everything resolved and volatility reverted back to a mean and the market cooled off and started climbing again, we saw what an expected return of 0.52% daily. This is the simplest example of past returns not being indicative of future performance. Moreover, this has absolutely no measure of risk attached to it. So, the year-to-date return is 22%. Roughly 22%. But we had to suffer this massive draw down for the first couple months of the year. That's not baked into this year-to-date return. So expected returns don't offer enough information for us to use them in isolation when evaluating some sort of trading strategy or investing strategy. Those are some very important concerns when it comes to the expected return for evaluating your investment or trading strategy.

Okay. Well, what about win rate? Win rate is the craziest thing I've ever seen cited in the trading space. On its own, it's an entirely useless measure. I can give you a strategy right now that has roughly a 100% win rate. It tells us nothing of the performance of the strategy. Tells us nothing of the strategy's wealth accumulated, amount of capital risked. It literally doesn't give us any information. And what I've done for you here is I've plotted three different trading strategies. one with a 100% win rate, one with a 99% win rate, and one with a 40% win rate. Look at the relative performance. The 99% win rate strategy had a return of 65%. The 100% win rate strategy only gained 2.97%. Meanwhile, the 40% win rate had a return of over 137%. This is why win rate is an entirely useless performance measure. We cannot assess anything of the overall strategy's performance. We know nothing about the risk. We know nothing about the return. It is a measure that can be used in addition to the law of total expectation to try to flip levers to enhance your strategy's edge. But that's not the context that we're talking about win rate in. We're talking about it as a performance measure in itself. And clearly looking at these three charts, it does not offer any reasonable information.

So how does this actually happen in practice? How can you have a 99% win rate, but then you have a crazy negative return? Usually, this is akin to selling insurance. So, imagine you're selling some sort of earthquake insurance, and you sell it on a month-to-month basis. Every month, you collect your premium, collect your premium, collect your premium, but then one month there's a crazy crazy earthquake. Now, you got to pay out all of that insurance. You had like a 99% win rate every month, you collected that premium, but then there was a massive earthquake and that took out all of your gains. This is the same thing as selling option contracts. You can sell option contracts, accumulate premium, accumulate premium, accumulate premium, but then at some point you're going to have to pay out that insurance and you're going to owe a whole bunch of shares. And that is exactly what happens here. You'll still have a very high win rate, but your overall return is going to get decimated by that insurance that you have to pay out. And if that isn't enough to convince you that the win rate in isolation is entirely useless, well, we also have this idea of forward-looking performance. Just because your strategy did have 100% win rate, it does not suggest that it is going to continue to have a 100% win rate. This 99% could go down to 95 to 90 to 85. We have no idea what's going to happen if you continue to trade this strategy. We do not know if its performance is going to remain stable, if it is going to decay. It largely depends on what it is you are trading and if there's anything to suggest that that particular strategy should be stable in a forward-looking sense. But again, this is the missing piece of the puzzle that we're trying to build towards.

What about generating cash? Why is that a useless performance measure? Well, $3,000 a month, right? I make $3,000 a month trading. Well, if you have a million-dollar portfolio, that's not good. You can just invest in US treasuries and earn the roughest proxy we have to a risk-free rate. As my old finance professor said, if the US is defaulting on its interest payments, then we have far more concerns than if it is a valid risk-free rate proxy or not. Nevertheless, let's take a look at the implications of this. Cash generated on a monthly basis. Okay. Well, let's just look at the bar charts first. It's like, hey, you know, my strategy generates $3,000 a month. Well, they're just buying treasuries with a million-dollar account, right? We're making roughly $3,000 a month. And you're not trading anything. You could say that you're trading treasuries, but you just locked your capital away at a a specific rate and you're receiving these payouts. This is not a trading strategy, but it is an easy way to cite the amount of money that you're generating every month versus what a trading strategy that might have some variability for whatever reason. Maybe you are actually assuming priced risk. Maybe this isn't some sort of true alpha and you don't have the stability that you're looking for. But nevertheless, this is the difference. This is why it's not a useful measure because we're just getting this versus maybe something like this. What we really want is something like this. This is a risk-free strategy. You're not trading anything. The amount of money that you generate in a on a monthly basis is not particularly useful especially if you don't know the overall portfolio size. Right? What we really want to see is portfolio size, relative risk and the overall growth over a certain period of time. That is like super super first step in evaluating a a trading or investing strategy. Those are like baseline, super baseline first step analysis of whether or not anything is is even remotely viable. We're not even talking about things like overfitting noise or invalid back tests or biased back tests and so on and so forth here. We're just suggesting like, hey, do we even have enough information in these performance metrics and whatever else it is we're analyzing to determine if there's any validity at all to an investing or trading strategy.

So, this is why cash is insufficient, right? If I showed you this or this, right? Doesn't this look kind of a lot better? It's a lot more stable and and yeah, that's great. But this isn't an impressive trading strategy. This is right. This generated 36% 36% return with very reasonable risk versus something like this where no risk was assumed at all and they just traded what? They just traded a treasury. So that's why cash on its own is an entirely useless measure. In any case, when it comes to expected return, win rate, monthly cash generated, there's no quantity you can give me. You can't tell me that you have a 100% win rate or a,000% expected returns or $10,000 a month generated in cash. That doesn't mean anything. It doesn't let me assess the viability of your strategy. When it comes to cash generated, I don't know what the size of your portfolio is. When it comes to win rate, I have no idea what your average winning trade or your average losing trade looks like. That dictates the profitability of your strategy. When it comes to expected return, I have no idea how many units of risk you had to take to generate that expected return. Which is why on their own, these performance metrics offer no information into the efficacy of your trading strategy, even in a backward-looking sense. We're not even talking about in a forward-looking sense and if it's tradable, but even in a backward-looking sense, insufficient information.

What about these what I'm calling less useless investing and trading performance measures? I thought things like the sharp ratio, sortino ratio, max draw down were the gold standard in evaluating trading strategies. Well, to understand why I'm calling these less useless and to understand why they are necessary but not sufficient to suggest you have a viable trading strategy, we need to understand this idea of volatility and risk. So, we've already talked about expected returns. We said, hey, we don't know what a reasonable benchmark is. It has no forward-looking predictability. What is this idea of risk? What is this idea of volatility? Well, it's the idea that if we have some sort of expectation, we can compute the expected squared deviations from the expectation. Okay, so what it is is a measure of dispersion around some sort of benchmark expectation. But we've already said that we don't have a fantastic measure, a perfect great good measure of expected returns. And now we're trying to measure expected deviations from that moving target. This is a tricky problem. So what we're trying to do here is get a sense of the relative risk to the return. But again, this is all in a backward-looking sense. This says nothing of what will be. It just says something of what was. However, we can bring in this idea now of a risk-adjusted return and that will help us assess the viability of a trading strategy. Tied in with this idea of stability, we can determine if it is in fact tradable.

To understand why these performance metrics are less useless, let's look at a couple examples and discuss the impact of this so-called risk-adjusted return that we are in fact computing. So let's look at these two strategies. We have strategy A and strategy B. Which one would you rather trade? Well, I would certainly rather trade strategy A because if I traded strategy B, then I would have this hefty draw down. I would have if I invested $100,000 I would have had a loss at some point of over $5,000 and I would have had to continue to trade through to recover and eventually reach the same performance that strategy A had. So if we just look in terms of the overall return of the strategy and the expected returns they are roughly the same. But if we account for risk using a sharp ratio or a sortino ratio or a max draw down then we can see that strategy A has what? It has a sharp ratio that is roughly comparable to strategy B. It has a higher sortino ratio and a lower max draw down. In other words, these performance metrics tell me, hey, strategy A didn't have all that much of a draw down. You don't have to sit through losses in order to make gains versus strategy B. You had to sit through a almost 12% loss. The Sortino ratio and the sharp ratio tell us the overall risk-adjusted return of the strategy. This is beyond just a normal expectation. This is telling us how much risk relative to the return that we earned are we actually getting by trading this strategy. Moreover, we're actually also removing the risk-free rate. The risk-free rate is something that you can earn just by investing in US treasuries, for example. So, if I can earn a guaranteed 3% by investing my capital in a treasury, then why would I create a strategy that has a 2% return? I would be losing out on 1% 2 minus three. So we're accounting for all of those things in these risk-adjusted measures. The difference, the only difference between the sharp and the sortino ratio is the sortino ratio doesn't penalize upside deviations. It's saying, hey, if we have a loss or a downside deviation like this strategy B, then we will penalize it more. That's why we see strategy B has a 1.77 sortino to strategy A's 2.09. So it will punish downside deviations but not upside deviations. So we can see they have comparable sharps. Sharp doesn't do this. It considers deviations on the upside and downside. But the Sortino, it only penalizes deviations to the downside. So when we're looking at these performance metrics, I call them less useless because we're getting a relatively full picture of the risk profile to the return profile of the strategy.

Though these performance measures give a more complete picture of the strategy, I still say show me the money. This isn't what I care about. You can come to me with the Sortina ratio of 2, three, five, whatever it is, but these are overrated. You will see people run back tests in so many different capacities. And I'm not just talking about overfitting. We'll talk about all the pitfalls in a moment. But these performance measures are overrated because if expected returns are not indicative of future returns and we're using that to measure volatility and risk, then that isn't indicative of future volatility and risk. And if we're putting those two together to come up with a sharp ratio, a sortino ratio, then those also aren't going to be indicative of performance in a forward-looking sense. That is why I call these performance metrics less useless. That is, on their own, they are a necessary component of a trading strategy. They give us a more complete picture of the relative return to risk profile. so on and so forth. But they alone are not sufficient to suggest you should trade in a live capacity.

I have a great example of why right here. Check out this trading strategy. So we're going to look at this one path. This one path had a sharp ratio of 3.1, a sortino of 5.6, a max draw down of 12%. That is a crazy performance profile. Okay. Well, what created this sample path? Maybe you are swinging at pitches. You're coming up with some sort of trade and you trade it. Maybe you have some sort of reasonable economic interpretation of your strategy and you you believe you have some edge in each trade. Maybe you are trading some sort of exposure that is performing particularly well for this time period. Maybe you even overfit your back test to noise and historic data. In any case, what actually matters here is the only thing that I care about when evaluating your trading strategy performance, and that is this idea of stability. Stability in a forward-looking sense. I'm not talking about chopping up your historic data into training and testing intervals. I'm saying if I was to deploy your strategy right now, would the returns that I observed in your back test, your performance profile, would that be reasonably stable if I continue to trade your strategy?

This is why when you look at people who are making discretionary trades, it's difficult to evaluate their performance because the stability of their performance metrics is what's in question. Yeah, maybe they have a great trade. Maybe they're using technical indicators. Maybe they're using the stars. Who knows? If they made money, it doesn't matter. But assessing the efficacy of their strategy comes down to stability. Is their performance stable over time? Does it degrade over a series of months, years, over a large quantity of trades? This is what I care about when it comes to your trading strategy. When you trade live, does your performance degrade or does it resemble the performance in a back test? Again, this is not chopping up your data in a back test into different intervals and then doing some sort of train test split, fitting it to training data, and then testing it out of sample. No, I'm talking about if I was to deploy it right now and I was to start trading your strategy. Do you have some sort of statistical mispricing? Are you bearing risk? What are you trading? And is it stable in a forward-looking sense? That is what I care about. These performance metrics are a piece of the puzzle. They are necessary. They are not sufficient to assess whether or not you should trade a strategy.

Here we have a picture that perfectly captures this idea. Maybe my back test went perfectly. Maybe I've been trading in a discretionary capacity for, you know, quite some time and I made a whole bunch of money and this is my performance profile and I think I'm the man and I know what I'm doing and I trade for the next couple years and look what happens. The out-of-sample performance of my future path if I keep trading that strategy whether or not it's a quant strategy a discretionary strategy whatever it is the performance degrades in a live setting so I only care about this idea of stability.

So let's take a look at a different example. I have strategy one that maintains the performance from the back test in a live trading setting. So what I have here is a back test sharp of 2.12 and a live sharp of 1.71. I can run reasonable statistics to confirm that yeah, this is in a reasonable range. This is stability that I'm looking for versus some performance that degrades in a live setting. Here I have a back test sharp of 2.31 in this example. And when I deploy the strategy live, I get a sharp of .51. That is not indicative of my back test performance. It's suggesting that hey there is a big problem in your back test. There is a big problem in the deployment of your strategy in a live capacity. The performance degradation could be due to a variety of different things. Maybe you were just trading an exposure that did well for a period and is now mean reverting. Maybe you had a lot of market exposure. Maybe the market is now in a contractionary cycle. Maybe you have a different exposure. Maybe you weren't creating any sort of statistical mispricing. What if you are? Well, maybe that statistical mispricing got crowded and it's no longer stable. That could suggest this performance degradation. Moreover, you could just overfit your back test. You can overfit it to historic data and then when you try to trade it live, you have terrible performance. These are just a few of the very many reasons why your performance at deployment could degrade severely relative to what your back test was. The stability is what's in question. And the necessary components, the good sharp ratios, the good sortino, the good max draw down, those are accessories. Those are things that you need to have in order to assess your strategy. They themselves alone aren't sufficient in evaluating whether or not you should trade it. In fact, look at this, right? You should not trade this. The performance has degraded versus something like this where it remains reasonably stable. This is precisely what I'm talking about. That is what you should look for when it comes to evaluating your strategy beyond these performance metrics. They are overrated. You will cite them in 10,001 back tests. You will see sometimes you may even observe some sort of strategy that's viable. You account for transaction costs, it's no longer viable. Maybe you even, you know, add a bias to your back test and the metrics that you observed are not indicative of the metrics that you're going to observe in a live trading setting because you polluted your back test with bias. There's 10 million in one reasons for this performance degradation. And to make matters even more difficult, if you did have a valid strategy and it was reasonably stable in a forward-looking sense, you had the performance metrics, the necessary performance metrics, they were stable, you very well could just have your strategy, your alpha crowded and it could decay. Now, you know, this could come back over time. This is a dynamic space. It's a time-variant space. Regimes change. Things change all the time. So, you know, these sorts of of out-of-sample performances, these these live performances that we see that resemble the back test, we don't expect them to persist indefinitely. They require very careful monitoring. This is not a set and forget. There's no such thing as a one-man set and forget hedge fund. I found an alpha. I'm just going to trade it. I'm going to lever up, trade it, and just leave it. It's not how it works. You have to constantly monitor for degradation and assess why it is. Are you actually trading a statistical mispricing? Do you have exposure? Are you bearing risk? All of these things matter when you assess your strategy.

Too long, didn't watch. Here is your executive summary. Performance measures like win rate, cash generated, expected returns, they can be misleading from the perspective of outcome. We do not know what good is for any of those measures on their own. We don't know anything about the relative performance of a trading strategy at all. In terms of expected return, risk measures like portfolio variance, volatility, and risk-adjusted returns certainly do give a more complete picture of your strategy, but are leaving out an important piece of the puzzle. This is a domino effect. Expected returns are not indicative of future performance. So neither is your risk measure volatility and variance. Neither is your risk-adjusted return. Now these all must necessarily perform well but alone are not sufficient to show that your strategy is tradable. Stability is in fact the only thing that I do care about. This is out of sample. This is in a forward-looking sense. It is the most important component of your analysis. Will you make money? Show me the money. Does the trading strategy actually offer a tradable mispricing? If it does, your alpha should persist. The performance that you observe in your back test should be reasonably stable in a forward-looking sense. This is suggesting something of your forward-looking performance. Of course, before some sort of crowding or regime change or so on and so forth, anything that may decay this performance. Now, without any evidence of stability, you can come to me with a sharp ratio of over 10, but that won't matter because you won't be able to execute a single trade in a live environment that resembles that performance profile.

Some future topics I would like to discuss technical videos and other discussions. We talk about randomness all the time. We talk about different ways to evaluate randomness and some sort of trading or investing strategy. We talk about the expectation variance, but is the market random? Is that really a fair assessment? Is it fair to use these tools and techniques? I'd also like to talk about the trouble with stationarity tests. This is certainly a more niche topic. Perhaps there's interest. Markov chains and hidden Markov models might have more of an appeal to a broader audience. And those are some of my my favorite models um in all of the theory of of stochastic processes. So I would love to discuss those. Um and of course our quant builds. I still want to build this earnings event options trading dashboard. I'd also maybe like to talk about these live regime switching models. Maybe we can implement the hidden Markov models. Um, and I've even toyed with this idea of building an automated delta neutral trading system for trading options in a a delta neutral capacity. Essentially constructing your delta hedge in an algorithmic way. That's going to do it for this video on why trading metrics are misleading. Unless, of course, you have some evidence of stability. I hope you enjoyed. I hope you learned something. If you like this video and you want to see more like it in the future, please comment, like, subscribe. It helps me out a lot. It is greatly appreciated. Join our Discord if you'd like to connect with other members of the Quan Guild community. Check out quantguild.com to master your quantitative skills. Other than that, thank you so much for watching and I will see you in the next video.