Transcription
Welcome to Moon Traders, the podcast where we explore the world of trading and learn from the best in the business. Join us as we deep dive into strategies, techniques, and the mindset of successful traders from around the globe. Our goal is to help you build your own edge in the market and achieve your trading dreams. So sit back, relax, and get ready to learn from the masters on Moon Traders.
In today's episode, I got to speak with Ernest Chan, who uses machine learning in a very unique way to gain an edge in trading. If you're interested in using machine learning or at least learning how machine learning is applied to trading, then you have to listen to this episode. Let's start the show.
Hey, Ernest, how are you today?
I'm doing well, thank you.
So, I've got a lot of great questions for you today. As I understand, you started as a machine learning engineer, but my question is, what persuaded you to go from researcher to trader?
Well, I think that my curiosity was piqued by the departures of a number of my colleagues at my research group at IBM Research. They went to a hedge fund, which was not very well known at that time, but it was very strange to us because we thought that our research group in machine learning was one of the top groups in the world. Why would they leave? It turns out that they made the right decisions, at least financially, because it turned out to be the most successful quantitative fund of all times.
So, you know, I got intrigued by their choice, and I also wanted to explore how machine learning can be used in asset management. That's why I joined the newly formed AI and data mining group at M Sany.
Very cool. And what was that firm that everybody left for?
It's called Renaissance Technologies.
Okay, I'm sure many of our listeners know of them. What does your day-to-day workday look like? Are you coding a lot, trading, researching?
Well, I have done less of that in recent years, as I manage an increasing number of people. So my job is becoming more like a product manager than a researcher. I basically define the objective of the project and the product, look for investors, look for quant team members, look for traders to implement them, and, you know, basically put the whole commercial system together so that it generates returns for all the stakeholders involved.
Very cool. What markets are you guys trading?
Well, our primary strategy trades stock index futures, but we also have strategies trading options. We have invested in traders that trade stocks, but I would say that the majority of the trading activities that we either invest in or engage in ourselves are in the futures and options on futures market.
Amazing. What are some of the top ways that you and your team use machine learning inside of your trading models?
Well, there are two ways. One is what we call corrective AI. That is to say, using machine learning to generate a probability of profit for each trading day. If the probability of profit is too low, we might decide not to trade that day or not to trade a particular strategy on that day.
The other one is what we call conditional portfolio optimization. So it is essentially a machine learning-based method to allocate capital to different strategies based on the market and the economic environment. Instead of saying that, oh, well, you know, we just invest equal amounts of capital in each of the strategies or we invest capital based on the inverse volatility, we update allocation regularly, maybe even every day, based on whether we expect the market will favor one strategy or the other using machine learning.
That's super interesting. I feel like a lot of people approach machine learning trying to predict price. Have you tried that in the past and just seen it doesn't work as well as these new methodologies?
Yes, I mean, that was sort of the naive application of machine learning. You know, everybody who started with finance and machine learning wanted to predict returns, but that turned out to be a huge challenge because the signal-to-noise ratio in that kind of prediction is very low.
So oftentimes, what the machine learning system learns are noise patterns that do not repeat themselves in the future, and oftentimes, it was not successful. At least I have never been successful in that particular endeavor. Maybe others have.
And, you know, I think that many experts in this field have also recognized that. One of the experts, Dr. Marcos L, used to give a seminar with a title that says, you know, the top seven reasons why most machine learning hedge funds fail. Many of the reasons are what I described. So we decided that we should apply machine learning in a somewhat different way. The two ways that I described are what we found to be actually successful.
Wow, that's super interesting. I think that's going to help a lot of people avoid trying to predict the price. You mentioned there's so much noise. How do you get your machine learning models to focus on the actual signal within all the noise?
Well, if we use machine learning for asset allocation or for corrective AI, it turns out that the level of the noise is decreased because they are filtered throughout the trading strategy.
So the trading strategies, you can think of them as a filter of noise because they use information that is just beyond the empirical observation of prices. They typically are built upon a causal theory of how the market works. For example, you might know something about the behavior of institutional traders, how they react to news, or you might have a theory about how options market makers work and how their behavior would affect the futures market, and so on and so forth.
So you typically have an in-depth fundamental understanding of some aspect of the market and have a causal model of how prices should behave that's beyond what you can just empirically extract from the market. With that kind of causal model, your strategy typically has a much more consistent behavior than the prices themselves.
Also, particularly, you might call it a signature risk exposure. For example, if your strategy is a momentum strategy, a strategy that longs volatility, it is very predictable that this strategy will be losing money when the volatility is low. You ask any momentum trader, they will tell you, okay, what are the circumstances that your momentum strategy will not perform well? They will say a very calm market, very low volatility market.
And if you have a mean-reversion strategy, the opposite is true. You would say that a tool environment where the volatility changes a lot—not necessarily high volatility, but where the volatility of volatility is high—is not favorable to mean-reversion strategies.
That's kind of common folklore for a lot of traders, especially quantitative traders. Now, machine learning can discover that, but machine learning can also combine multiple variables, not just volatility. Maybe they look at, you know, maybe your strategy trades only oil futures, so it would not necessarily use equity volatility as a predictor of where the oil futures strategy would make money. It might look at the oil volatility and so on and so forth.
So the pattern is, in other words, the signal-to-noise ratio in detecting whether a strategy is possible is much higher than the signal-to-noise ratio of predicting the direction of the underlying market itself.
Interesting. So it sounds like you're building strategies and then having your machine learning kind of select which one to turn on and off at each time. What is your typical strategy that you feed in? What time frame are you trading on, and what's your average hold time for a trade?
Well, that obviously depends on the strategy, but one of our perhaps strategies with the longest track record is a strategy called Tail Ripper. It is a quantitative strategy that is implemented using stock index futures. It is an intraday strategy; we never hold positions after the market close, and the typical holding period can range from seconds to hours.
So, but it never exceeds one trading day. Of course, there are other strategies and other traders in our fund that take longer positions. We try to cover a whole range of different frequencies so that we get some diversification there.
Yeah, that makes sense. What would you say the percentage is intraday versus, let's say, multiple days?
Well, I would say that our core strategy takes up maybe about one-fifth of our buying power, and the strategies that hold longer-term take out the other four-fifths. But, you know, many of them are sub-strategies. So our core competence really is still in intraday trading, not taking overnight positions.
Very cool. Now, that kind of touches on risk a little bit. How do you manage your risk, and what percentage of capital do you risk per trade?
Well, one way is by what I mentioned as corrective AI. So based on the probability of profit, if the probability is low, we decrease the amount of risk we take. And of course, there's a maximum amount of risk as well.
But the maximum amount of risk, you know, we don't really adjust very often. In terms of this maximum, we have a stop loss, obviously, because it's a momentum strategy. You know, stop loss is appropriate for momentum strategies, so we do have a stop loss set as a percentage of our capital.
But other than that, we don't change the maximum risk. Of course, we are always looking for alternative implementations of the strategy that would lower the risk, and, you know, sometimes we would trade differently, looking for trading different instruments, different time frames, different overlays, and so forth.
Ernest, this is fascinating. How often do you retrain your models?
You know, a machine learning model doesn't have to be retrained too often because, particularly because of corrective AI, I generally feel that if you have about 10% of your data as new, you should retrain the model.
So for example, if you have, you know, 10 years of data, you might not need to retrain your model more than once a year and so forth.
That makes sense. Speaking of data, how much data do you think is needed from your experience to train a great model?
Well, I would say, generally speaking, at least 500 data points. So if you have a daily trading model, you need at least two years to train it. But, you know, preferably a thousand points would be more ideal, so four to five years of training data I think is ideal.
But in no circumstance should you train it with less than two years of data. But also, sometimes it's not just a matter of length; you want the training data to capture different regimes of the market.
So for example, you know, you want to capture regimes where the market volatility is very low, like present time or in 2019, but you also want to capture data regimes which are extremely volatile, like 2020 or 2021 or 2022.
So, you know, you want to have a mixture of different market environments, increasing interest rates, decreasing interest rates, and so forth. Given that kind of requirement, you would typically find that, you know, six years is a good number to capture the different sorts of circumstances.
That makes sense. So you mentioned you input strategies, and then your ML kind of looks through them. But I've also done a little research, and you once said that you could build a machine learning model that could output profitable strategies. Maybe not like a five Sharpe or anything crazy good like this, but they can be good profitable strategies. How do you go about that?
Well, I mean, with the latest events in machine learning, it is conceivable that you can use a large language model to learn from all the published strategies out there. You know, after all, there are many blogs, many books, many papers that describe different strategies, and you can use a large language model to digest them.
If you specify a particular kind of risk profile, they may be able to regurgitate or recombine those elements that they have learned from these papers into a strategy. I have not seen anybody successfully publish any work that—well, maybe they won't publish it if they succeed—but, you know, this is one area that I think it's not necessarily that the machine learning can come up with new strategies.
It is that it can combine elements of strategies that other people have discovered and then put them together based on your specification.
What are the main features you're inputting into your machine learning models?
Well, features should be as many as possible. So, you know, there's no such thing as too many inputs. In traditional quantitative finance, there's a fear of overfitting. You don't want to have too complicated a model; you don't want to have too many parameters.
But because we are using machine learning not to generate predictions on prices but on predicting whether a particular strategy is profitable or allocating what amount of capital to a particular model, the fear of overfitting is reduced.
So generally speaking, you want to have as many input features as possible, and that would span the universe of price-driven variables.
Let's take a quick break from the episode. So, because if you're looking at a chart right now and stressing over a position, that means you're trading with emotion, and emotions trump your logical brain. So you're not going to be as profitable if you're trading by hand.
Now, you might be a unicorn, but for the rest of us, I would definitely learn how to algo trade, and you can learn that in the Algo Trade Camp. Just go to agotradecamp.com. Automate all your trading strategies, get rid of emotion, and get back to living life.
You know that prices would include any markets—prices that are equities, fixed income, futures, forex, options markets, indices would all be considered prices.
So these are—and also you want to include fundamental data from various stocks, bonds, and so forth, or companies. You want to include macroeconomic variables, you know, and so forth.
So it's really—and, you know, obviously nowadays, with, again, with large language models, you want to include sentiment. In other words, information that is extracted from unstructured data. Unstructured data can include certain, most commonly, text or even audio feeds, videos, and furthermore, you might want to include data extracted from English data and the like, or any other kind of sensors, weather, and whatnot.
So all of the data nowadays can be converted into numerical form and be fed into a machine learning program. And they should, you know, if you can afford it, get as much as possible because, after all, many of the biggest hedge funds already are doing it. They can afford it, so certainly they're doing that.
That makes sense. You touched on this earlier, but I think it'd be helpful for the audience to understand a little bit better. Can you tell me what your thought on corrective AI is and your approach with corrective AI?
Well, it's very simple, actually. You know, let's say you have a trading strategy, and you had a track record of some daily returns. You wanted to build a machine learning model to predict whether the next return is positive or negative, and with what probability that it is positive.
To train this machine learning model, you need the past daily returns as an input, but you also need as many predictors as you can find, or what people call features in the machine learning world.
So again, these features would include everything from price-driven features to economic and whatever you can get your hands on. You train this machine learning model using these features to predict the sign of the next day's return with a probability. That is corrective AI because once you have that probability, you can use that to determine how much risk you want to take. Maybe the amount of risk you want to take is zero.
I love this approach, man. It's really great. What are some good non-price inputs or data feeds that you know provide good results for your machine learning?
Well, I mean, news sentiment is usually the first thing that comes to mind. A lot of people would apply large language models to convert news headlines into sentiment—news headlines on stocks or news headlines on the economy or anything financially related.
So, you know, you can also use it to scan through all the company press releases or quarterly reports, transcripts of earnings calls, and so forth. Any—or even research reports by analysts—these can all be converted into a numerical score, 0 to 1, that indicates the sentiment on a particular stock or anything else, or on the direction of the US dollar index or on the futures of particular futures, you know, oil futures or corn futures or whatnot.
So, yeah, I mean, that is what is generally called unstructured data. It is what is contained in reports that are written in text.
Very cool. So let's pretend you're just starting over again, but you have all the knowledge you have now, and I sent you 10 strategies and said, "Ernest, what are you going to do with these 10 strategies? They're 1.2 Sharpe, all 10 of them." What would you do from here with your machine learning models?
Well, we would typically apply our conditional portfolio optimization technique. So we would use these daily returns of all these 10 strategies as an input and then again apply our set of pre-engineered features to learn to find out historically under what kind of regime we should allocate more to strategy one versus strategy two and so forth, and how much.
Now, by regime, people often think of this kind of discrete category: "Oh, we are in a bull market," "Oh, we are in a bear market," or "We are risk-off, risk-on." But by regime, I'm thinking really in a high-dimensional space. To me, a regime is not defined by on and off or these kinds of discrete categories; it is a point in a high-dimensional space.
This high-dimensional space might have 20 axes, maybe 50 axes, maybe 100 axes. How many axes is determined by machine learning; it's determined by training a machine learning model. A regime is just a point within this high-dimensional space.
So with that point of view, we don't have a sudden jump from one regime to another. To me, it is a smooth transition. You know, we don't believe that, you know, unless you have a sudden news event, such as maybe a nuclear detonation somewhere or some major default from a major company, other than those extreme events, typically the regime changes smoothly from one point to another.
It is possible that it can make a jump from one point to another in this high-dimensional space, but generally speaking, most of the time, it is a smooth and continuous adaptation of capital aggregation from one day to another.
Interesting. And how do you measure success? Are you looking at Sharpe ratio, and if so, what Sharpe are you looking for to allocate to that model?
Yeah, by default, optimization typically—the objective is the Sharpe ratio of a portfolio. After all, that is what the modern portfolio optimization method will recommend. They would recommend, you know, for example, the tangency portfolio is the one that has the highest Sharpe ratio.
But that doesn't have to be the case. Some portfolio managers prefer a portfolio to have the smallest expected drawdown, or expected shortfall, or conditional value at risk. We can use that objective. Others want to have a minimum drawdown; we can also use that as the objective function.
For me personally, when my trading strategies are concerned, the Sharpe ratio is not the most important. What is most important to me is what people call the Calmar ratio or the M ratio, which is the compounded annual growth rate divided by the maximum drawdown over some look-back period. That is my preferred metric for strategy, but it's really up to the portfolio manager to decide.
Absolutely. What Calmar ratio are you looking for?
Well, you know, of course, the minimum we want to have is one, but if you can get up to two, you are in a very happy place. If you can get to three, you are really one of the top 1% of traders.
Really amazing. Ernest, I've heard you talk about long-tail events, and you could have an algo that just runs once every two years. Can you talk to me a little bit more about long-tail events, what they are, and just explain it for the audience?
Yes, there are many names for this kind of strategy. One is long volatility, meaning you are buying volatility; you benefit from increasing volatility. Another name is called a tail hedge strategy, meaning that this strategy will generate returns when there are extreme moves in the market.
Some people call it a crisis alpha strategy because it only generates alpha when there's a crisis, which is essentially an extreme event, extreme movement in the market. Finally, some people call it a convexity strategy because if you think of a convex strategy like a U-shaped strategy, the equity in that strategy will sharply go up as you move further from the mean.
Whereas if the prices remain close to the mean, it's very flat; you don't generate much return. But the return that they generate sharply accelerates as you go farther from the mean. That's what people call the convex strategy, as opposed to a concave strategy where, you know, when market prices move away from the equilibrium, you suffer big losses.
That typically is what happens when you're short options, right? If you're short options, you have a concave strategy because ideally, the price doesn't move, and your options expire worthless. But if the market moves significantly, you'll be deeply in the money, and your position can generate unlimited losses.
So long volatility strategy is the opposite of the kind of strategy that many option traders favor, which is shorting volatility.
How do you go ahead and backtest your models to build confidence?
Well, there are many standard backtest platforms out there, but generally speaking, we would just backtest it using standard programming languages, whether it's Python, R, MATLAB, or even C, C++, Rust, whatever programming language that you're comfortable with.
Generally speaking, there are two ways of backtesting. One is using one of those so-called array languages, such as Python, MATLAB, and R. You backtest using an entire array at the same time. So essentially, you have an array containing all the prices, and you generate trades on all the prices all the day simultaneously and backtest it in an array-oriented way.
It's very efficient; it's very simple, but it's also very much prone to look-ahead bias. The other way is a more traditional loop-based backtest where you feed the trading logic one bar at a time and generate signals and then compute returns for the next day.
That looks very old-fashioned, very inefficient, and certainly not a very modern approach, but it minimizes the possibility of look-ahead bias.
The key, however, to any backtest is to avoid look-ahead bias. If you, you know, the more complicated the strategy, particularly if it's a machine learning-based strategy, the more likely you will insert look-ahead bias, and the result will look wonderful, but it is not replicable in live trading.
That's another consideration in backtesting: one really should not be using backtests as a way to discover new strategies because if that's the case, you will always eventually find a new strategy just by backtesting enough.
A backtest can be thought of as an experiment on historical data. You can't perform an experiment like in physics where you change the condition of the market; you can't really change the condition of the market. All you do is use existing historical prices.
Nevertheless, that backtest can be used to reject or confirm a particular theory. For example, you believe that every time an earnings announcement exceeds the expected earnings by 5%, the market is going to go up overnight. You should get in after hours.
So let's say that's a hypothesis, a theory that you have. Well, it's easy to confirm or reject that. But, you know, if you don't have that theory, you would have to try, you know, maybe 100 different entry and exit thresholds based on an earnings event, and eventually one of them will work.
But the fact that it worked historically doesn't mean that it will work in the future because this is really without a causal model. You are doing a blind search, and it is very often that you get a selection bias.
So to avoid selection bias is another important pitfall to avoid in backtesting. The first pitfall to avoid is look-ahead bias. That is, you know, it's fairly easy to avoid if you know how. The second one is to avoid selection bias or data mining bias, or whatever you call it. That is more difficult; it's more subtle.
Can you talk to me a little bit more about what look-ahead bias really is?
Well, I mean, a very simple but subtle look-ahead bias is that, for example, some people wanted to do pair trading. In pair trading, let's say you want to trade Microsoft against Google. The question is, well, how many shares of Google versus one share of Microsoft?
The typical way to find that out is to run a linear regression between, let's say, the prices of Microsoft and Google. But in a trading strategy, you should only be able to use the past prices to run this regression.
But a lot of traders forget that and they use all the data that they have between Google and Microsoft to determine the ideal hedge ratio and then use that hedge ratio to feed into their pair trading strategy, maybe buy low, sell high the spread.
Well, you know, that has already introduced look-ahead bias because you wouldn't know that ideal hedge ratio unless you have all the data, including future data used for the regression.
So a proper backtest would use a rolling linear regression. The last data point would be, for example, today's price, not tomorrow's price, not the price the day after, in the historical context.
So look-ahead bias can be very subtle. You know, because of course, nobody will be silly enough to say, "Oh, I'm going to deliberately use tomorrow's price to decide whether I would enter today." Nobody would be so foolish as to do that.
But there are many other subtle ways that you can seep into your strategy, such as in determining the hedge ratio of a pair.
Do you find your models select more mean-reversion strategies or momentum-type strategies or some other?
It's not a question of selection. Obviously, ideally, you would want to create a balance of the two types of strategies because, like I mentioned, momentum strategies are long volatility; they are convex. They oftentimes have a negative beta relative to the market, whereas mean-reversion strategies are short volatility; they are concave, and they oftentimes have a positive beta to the market index.
Let's take a quick break from the episode real quick to talk about automating all of your trading strategies. There are so many different tickers and symbols and opportunities every single day. It's impossible to compute it all as a human.
So why don't you have a computer trading for you? I don't know. I'll teach you how to algo trade at algotradecamp.com, step-by-step tutorials. I wasn't a coder before; you don't need to be a mathematician. You don't need to go to school for this. You're already ambitious enough to trade.
Let's just automate that strategy so we can remove the emotion out of it.
So you want to have the two running side by side or in some balanced mixture in your portfolio so that you would have a market-neutral mix, right? Zero beta.
Secondly, if you think of their exposure to volatility, you know, again, the obviously the mean-reversion strategy will have a negative loading to volatility, whereas momentum strategies have a positive loading, a positive exposure.
So again, with respect to volatility risk, you also want to balance the two types of strategies so that you are market-neutral with respect to market movement and also neutral with respect to the volatility risk and so forth.
So it is not that, you know, we randomly find that we find more mean-reversion strategies than momentum strategies. We try to balance them both.
One other comment I want to make is that oftentimes there is an illusion that it is easier to find strategies. Many people are able to find new mean-reversion strategies because they forgot to include bid-ask spreads in the transaction cost estimate.
If you did not include sufficient transaction costs in your backtest, many mean-reversion strategies will appear profitable, but they won't be in real life because you have to pay the spread, whereas momentum strategies do not suffer from that.
So it's actually a bit harder to find momentum strategies than mean-reversion strategies, but that's not the reason that you should trade more mean-reversion strategies.
Interesting. So with machine learning, you know, opposed to trying to predict price, you could predict sales, company earnings, market cap predictions, all that good stuff. In your opinion, what would be the best way out of those couple examples to apply machine learning and why?
Now, that is a very good point. We earlier said that machine learning had a very difficult time in generating alpha because of the low signal-to-noise ratio.
The reason for that low signal-to-noise ratio, in turn, is because of arbitrage. If the machine learning program can predict, "Oh, you know, Google always goes up when it has good earnings," well, everybody—not only your machine learning program, but maybe one million other ones and also human traders—would observe the same, and that phenomenon will go away gradually.
Maybe not instantaneously, but after a few months, a few years, it will go away because more and more people will recognize that it is not possible that a strong signal will be only noticed by you.
It's not possible. The only way the machine learning can be useful is it can detect some subtle pattern that no one else can detect. But how would you know that that subtle pattern is not noise? It's quite difficult, right?
A weak signal could be because it is noise, or it could be because no one else has discovered it, and you are the smartest and the greatest machine learning program. You detected it. It's hard to differentiate the two, quite frankly, because of arbitrage.
This is the only kind of signal that your machine learning program can find. On the other hand, if you are using machine learning to predict quantities that cannot be arbitraged away, for example, as you suggested, predicting earnings, predicting cash flow, predicting any sort of financial statement quantities, or predicting whether the Fed is going to increase interest rates or predicting whether the non-farm payroll number will surprise or not, well, those quantities cannot be arbitraged away.
You know, the US Labor Department is not going to change their non-farm payroll number just because somebody else was able to predict it, right? You know, those quantities are not subject to arbitrage, and so machine learning has a much easier time predicting them with a lot of accuracy.
But the problem is, okay, so what if you can predict them precisely? I mean, you can't predict them with 100% accuracy anyway, but maybe 60%, 70% accuracy. But even with that level of accuracy, can it be translated into a trade? That is an open question because it may be that you can predict non-farm payroll quite accurately, but the reaction of other traders to that non-farm payroll number is much more unpredictable.
Also, just because you can predict it with 70% accuracy, maybe a lot of economists can also predict it with 70% accuracy. So maybe you can predict it with maybe 5% more accuracy than the consensus.
So what is that 5% accuracy increase enough to generate arbitrage profit? Not necessarily. So, you know, that's the dilemma in applying machine learning is that the quantities that you can predict with high accuracy do not necessarily translate to good returns.
Wow, Ernest, this is really, really got me thinking. For a new researcher that wants to approach machine learning for trading in the methodology you're using, where we're selecting our strategies instead of trying to predict price, what are some good models that a new researcher should look at? You know, LSTM, neural networks, or yeah, what models should we be looking at?
You know, I generally speaking, you know, again, my use of machine learning typically is more indirect. So I'm not going to suggest a particular network architecture. Oh, you should use a transformer-based neural network, or oh no, you should use gradient-boosted decision trees.
You know, I have some opinions, but those opinions regard what machine learning model, what machine learning algorithm should be used for certain data, not for the objective of predicting prices.
What I feel that one should learn from predicting prices is actually much more classical. From all the profitable strategies that I've seen myself, they typically, again, have a causal component or derive from some fundamental observation of the market.
So for example, the fact that there is a lot of—whether you are an options market maker or whether you are a market maker for volatility futures or variance swaps and so forth—you need to balance your book by the end of the day. You want to remain delta neutral, or you want to remain neutral with respect to all these other Greeks, and that induces certain trades.
So if you know the market has a strong positive delta, you might find that there's a positive momentum in the market towards the end of the day.
So these are the sorts of observations based on an understanding of market microstructure and what different players, particularly what different institutional players are doing.
Now, it may be less easy to predict what retail traders are doing. Maybe, you know, because retail traders may be affected by fads or Reddit or whatever latest fashion that there is.
So it's maybe less predictable, but you know, certain institutional players, their behavior can be more predictable. So based on that kind of analysis, we build a quantitative trading strategy to capture that.
So the basic idea is actually usually quite explainable, quite understandable. This is oftentimes, you know, the fundamental phenomenon is often documented in books on options, books on futures, books on forex markets, books on, you know, the agricultural market or whatnot.
This is what, you know, one would call people who have fundamental knowledge or discretionary knowledge would do, except that we as quants would take that knowledge and quantify it, abstract it, simplify it, and quantify it.
That is very different from saying that I should use an LSTM to make predictions on prices or I should use instead some CNN to predict prices and whatnot.
So I find it much more fruitful to understand the market from a fundamental point of view than from a purely what is called technical point of view.
You mentioned before in the past that you do trade a little bit of crypto. Are you still trading crypto at all?
Not so much at this point.
I was curious because you had mentioned that there are some strategies that still work in crypto today that haven't worked for the last 10 years in equities, and I was wondering if you had a couple of examples of those.
Yeah, I think that, you know, I have heard a lot of the statistical arbitrage strategies that people have retired from the equities market continue to work in crypto. There are certainly people who reported a high Sharpe ratio. We, for a while, have been able to generate that as well.
But just like the equity market or any other mature market, the crypto market is also subject to regime shifts. Oftentimes, the strategies that work well in one regime would cease to work in another.
So, you know, even though it might seem easy to have these traditional techniques that work, that doesn't guarantee that it will work forever, and that's what we find as well. So that's why we stopped trading one of the strategies that we observed has kind of underperformed in a different regime.
Do you use any grammatical evolution in your trading, and if so, can you explain that a little bit?
I'm sorry, what evolution?
Like grammatical evolution.
No, we actually don't use that.
Oh, that's okay, cool. What are your thoughts on LLMs? I know you touched on it earlier. How do you see that being used in the future now that they're kind of being rolled out?
Well, I mean, the most obvious use of them is to generate features for other machine learning programs. For example, as I said, to convert all the text-based, in general, any unstructured data into numerical inputs to other machine learning programs for portfolio optimization, risk management, or even alpha generation.
So you can convert text, you can convert speeches, videos, images, what have you, into numerical data, and that's one obvious use for, you know, the—and other similar generative AI techniques.
Now, I should say deep learning techniques. The more advanced use is what I also mentioned earlier, which is that you can use LLMs to learn from all the existing literature about how one can trade in a certain market.
For example, there are all these books about options trading, you know, long for short for calendar spreads, this and that, and you say, "Well, I have this—I want a strategy that, you know, has this risk-return profile and so forth. What strategies would you recommend, given that you have learned from all these books?"
Well, that's something that, you know, at some point should be able to provide, even though it may not be able to right now. But it's something that they can suggest, and they may hallucinate, and you don't have to trust it, but you can easily backtest.
Right? You can ask an LLM to, "Hey, suggest 10 trading strategies that are long volatility in the forex options market," and it can go out and read all these books on FX options and volatility trading and whatnot, and it will spit out—hopefully, it will understand those papers and spit out 10 strategies.
Maybe nine of them are nonsense; the LLM doesn't really understand what it's reading. But maybe one of those strategies actually makes sense, and you can go ahead and backtest it, and it generates good returns.
That would be a good way to use LLMs to generate trading strategies. Now, that's a theory. I haven't seen any published work that shows that it is working.
In between the two, there is a possibility that you can use LLMs to directly convert sentiment to trade, and that's actually the project that we are working on.
So we are working on a project that uses the Federal Reserve chair's speeches as an input on a high-frequency basis, every sentence, for example, and convert that—every sentence that he says in a press conference—into a sentiment score and trade based on that sentiment score.
So it doesn't go through another machine learning program; it just uses this sentiment to trade on that, and we have achieved some promising results that we are actually going to report on in our generative AI workshop in October.
So that's a project that we are actually actively working on, and it's showing some promise. So that is in between the two usages, right? One usage I said is to just simply generate features from unstructured data and use it as an input to other trading programs.
The second one is to completely generate new strategies from reading it. This particular application of generating trades based on sentiment is sort of in between the two, and that's actually what we're working on.
And you might call it, you know, it's actually the first project of generating features is already done. It's well-known; people actually have been doing it for years.
The second project, which is using LLMs to generate brand new strategies from reading, has not been realized yet. Whereas what we are working on is in between. It is showing promise, not yet, but hopefully, we'll find the results and report it in our workshop in a few months.
I look forward to that workshop. Ernest, this has been a fascinating conversation. My mind is going crazy. In closure here, can you tell me a little bit about your three books, Predict Now, and also where can listeners learn more and work with you?
Yeah, the three books—Quantitative Trading, Algorithmic Trading, and Machine Trading. Now, it's very interesting. If you ask ChatGPT who are the best algorithmic trading book authors or what are the top 10 best algorithmic trading books, oftentimes my three books will show up in the top 10.
I don't know if it's hallucinated, so take it with a grain of salt. They do so. Predict Now is the company I started three years ago to offer this kind of machine learning assistance to traders.
We don't want to use machine learning to tell them what to trade; that's the job of the trader. We use machine learning to increase their Sharpe ratio, to decrease their risk using corrective AI, using conditional portfolio optimization. That's the service we offer through Predict Now.
We also offer, like I said, a workshop on how to use large language models to convert unstructured data to tradable signals. That's what I was just talking about, so we offer the workshop through Predict Now as well in October.
Awesome, man. I will definitely be there. Ernest, thank you so much for your time today. This has been an amazing conversation. I hope you have a wonderful day.
Thank you for inviting me. Take care.
All right, that's a wrap. If you enjoyed that episode as much as I did, definitely jump into my free Discord for traders. We have the biggest trading and algo trading community out there, and it's free to join.
All you have to do is search on YouTube, "Moev," find one of my videos, and you'll see in the description your free invite to our Discord. So I look forward to seeing you inside the community and on the next episode. Have a wonderful day!