transcribe

EDHEC Speaker Series | “Artificial Intelligence pricing and the Virtue of Complexity” by Bryan Kelly

EDHEC Business School · 59m · transcribed Jun 2026
More from EDHEC Business School Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Transcript

0:00 [Music] Hello everyone and welcome to the season five of the edex speaker series the future of finance. I'm Emanuel Joenko. I'm the executive director of the guad finance program here at EDC and it's my pleasure to kick off this new season with uh professor Brian Kelly. So well I don't have to to present to Brian that mean I think that everybody knows Brian. Brian is professor of finance at Yale school of management. He's the head of machine learning at Aqua Capital. Uh he's been very very prolific. Uh his research spans asset pricing financial econometrics and and the topic of today is a super interesting virtue of complexity. Uh so I think that today Brian will will do a focus on how complexity or highly parameterized factor models could be of interest for cross-sectional pricing. I want also like to thanks my colleagues uh professor Raman who is also well known I don't have to introduce himself uh is our expert on asset pricing too and I'm very pleased to have today one of our bright student from the MSE in financial engineering Tomma Ren. So without further ado, uh Brian, the floor is yours.

1:20 >> Great. Thank you very much. It's a pleasure and an honor. Um hopefully my slides are visible now. >> Yep. So um I'm going to be presenting on this topic of virtue of complexity which is um a sequence of papers that I've written uh with my longtime co-author Saman Malamode as well as with a very various other co-authors um a combination of our students at Yale and EPFL. Um eventually I want to work towards discussion of the role of complexity in asset pricing models in particular in cross-sectional asset pricing models. Uh but to get to get there I want to spend a little time first talking about the precursor paper uh the virtue of complexity and return prediction to give us a little bit of a a foundation before moving on to the more complicated environment. Um but I I want to start off with a bit of an overview of the role of complexity in financial modeling. Um so I've put a quote on the slide here that's rooted in this idea of a principle of parsimony which is an old idea in statistics. It goes at least back to Tuki. Um, and I put this quote up because this is one of the textbooks that I was trained on when I learned time series econometrics, a famous uh textbook by George Box, a very well-known statistician. Um, and in the first chapter of this textbook, it's called something like rules for effective model building. There are four rules and rule number one and chapter number one is this. It is important in practice that we employ the smallest possible number of parameters for adequate representations.

3:02 The reason why I start here is because I want us to appreciate that we were in some sense taught a bias in favor of small models. And this bias clashes to a large extent with the world of modern statistics and the success of modern prediction models in lots of different applications. Um, of course, many of us will be familiar with the role of large language models. You take a model like GPT, which has over a trillion parameters in it. Even in return prediction analyses, right, asset pricing modeling. Take for example some of my earlier work with Deng Shu where we studied neural network models for the cross-section of returns. And in those models, they're puny by sort of Silicon Valley standards, but they still have 30,000 or so parameters in them.

3:51 These are models that have worked pretty extraordinarily out of sample compared to a lot of the simpler models that we have at our disposal. I like to start here because if you look at what's happened in machine learning from the perspective of a box Jenkins econometrician, that seems ludicrous, right? Right? It seems that these models with 30,000 parameters, let alone a trillion parameters, going to be massively overfit. And if we think they're massively overfit, then they're likely to behave terribly when we go out of sample. But this is of course not the case. We use every day extremely heavily parameterized models, whether we're using voice recognition software on our phone, whether we're using chat GPT to look up the answer to any question that we might be interested in. We are used to using models that are demonstrating on a day-to-day basis their success out of sample. Right? So I just want us to be cognizant of the fact that there is an enormous expansion of models that are massively massively heavily parameterized. So heavily parameterized that they very easily fit the training data yet they perform exceedingly well out of sample. And it's often the case in fact that if I have two models that both exactly fit the training data, it's not uncommon to see the larger of those two models work better out of sample.

5:14 Right? And this is not just something that's happening in image recognition or textual analysis. It happens in finance too. Right? So a lot of the motivation for this line of research is to recognize that well the work that I and other people in our field have done have documented in my view unambiguously clear empirical gains to using machine learning models and finance applications. Right? These are not small gains. The typical gains from using a machine learning formulation to say represent an SDF or represent a return prediction model are oftentimes more than a few percent. They might be a couple times better than comparable simple models. Right now I think one of the difficulties in working in this literature is that we don't have as much theoretical understanding of the properties of these heavily parameterized models. Right? And without without a good theoretical understanding, it's natural for there to be a lot of skepticism about the reliability of models like this. All right? So you can think about this as really being the genesis of this research agenda. What I'm hoping to do here is build the case for why we should be using machine learning and finance.

6:31 And I think the very best way to do that is with development of the theoretical properties of highly parameterized models. Okay. So that's all we do in this literature. And the main theoretical result that we've shown in a couple different environments now and I'll take you through those is that model performance measured in a variety of ways. You can measure it in terms of statistical performance. You can measure it in terms of economic performance say in utility terms or sharp ratio. Model performance of asset pricing models is generally increasing and the number of parameters that those models have. So what I'll try and do is spend a little bit of time explaining the intuition.

7:09 This is a short presentation, so we aren't going to be able to go into too much detail. And I'll also show show you some empirical evidence that supports uh the theoretical predictions. So with that said, let's get directly into the environment. I want to give you a stylized environment to begin with to get a little bit of our feet underneath us thinking about heavily parameterized models. So in the stylized environment, I want you to be thinking about a problem where we're trying to forecast a single time series of returns. So think about R here as being the return on the aggregate market portfolio. And I want you to take as given just for a simplification to make this concrete that we know the set of predictors.

7:49 Let's call those G. So G will be a vector of predictors. Maybe they'll be sort of you know the usual suspects that we like to think about in macro finance for predicting the equity risk premium. So when I think about what finance and economic theory is good at and what it's less good at, I usually think about economic theories being pretty good in terms of giving us qualitative predictions about the way the economy behaves. So things like, oh, here is a set of variables that might be useful for understanding future return distributions.

8:22 Where I think economic theory is often less clear is on the specific details of functional forms. So if I think about a predictive function f here that relates the current predictors to future returns. You know often times we make a simplifying assumption that f is linear. Often times our theories make predictions that f is linear but that is really a choice on the part of the researcher a choice on the part of the theorist for the most part. So I want to approach this prediction problem from this concrete setting where I posit that I know the existing predictors G but I don't know the prediction function F.

9:04 All right. This is really the foundation of thinking about machine learning. It is really just the foundation of nonparametric econometrics. We know that if we want to approximate an unknown prediction function f an effective way to do it is with basis expansions. We have many so-called universal approximation theorems that allow us to approximate an unknown function f to an arbitrary degree of accuracy as long as we include enough basis terms of the original arguments of that prediction function. Right? So each S here and S is going to be indexed from one to P where P is the number of basis terms in my expansion is just some nonlinear transformation of the original predictors. So I add them all up in some way with some weight vector and I transform them through what machine learners would call an activation function, but it's just a nonlinear basis transformation. Maybe a max operator, maybe a sign function or so on. And what the basic universal approximation theorem tells us is that as I make the number of terms in this approximation large in the limit, I'm able to recover the original unknown prediction function.

10:20 Now to combine all of these basis terms, I need to estimate some parameters, these betas here. All right? And this approximation holds in theory. I still have to estimate this approximation in practice. Right? So what I've just described is of course the underpinnings for a neural network model in a neural network. I begin with some fixed input predictor variables called them G. And then what I do is I convert those G's into a bunch of basis terms. We call those neurons in a hidden layer. All right? So those neurons will correspond to my variables S that I've constructed out of my original raw G data. Once I have these G's which are expressing the nonlinear attributes of my predictors, I'm going to recombine those basis terms in a final linear prediction step.

11:10 That's the estimation of these beta coefficients. Right? So the general pipeline when we're thinking about a machine learning prediction problem is that we posit some unknown prediction relationship and then we approximate that unknown prediction relationship with a linear function that uses a nonlinear transformation of each of the predictors. All right. So the question is how big of a model should I use? How many of these basis terms P should I use? Because there's a clear tradeoff, right? The approximation theory says I should make P very large. If I make P very large, then I have a good approximator. My specification error gets smaller and smaller as P gets bigger and bigger. But as I add all of these terms, well, now I have more and more parameters to estimate. And we know as we estimate more and more parameters holding the amount of data fixed, the estimation of those parameters is going to induce sampling variation. And that sampling variation can get worse with the size of the model. Right? So, that's the trade-off we're thinking about. And the central research question is which P should the analyst opt for and we're going to put forth an answer that might seem a little bit counterintuitive at first, but hopefully we'll be able to understand why this happens. And the answer is that you should use the largest possible P that you can compute.

12:32 adding more coefficients, adding more nonlinear terms to your approximating model dominates the cost, the statistical cost, the variance costs of having to estimate those additional predictors. All right, so again, in the interest of time, I'm not going to be able to go through this in too much detail. Um but I what I want us to appreciate is that when we think about the number of parameters in our data set sorry number of parameters in our model we need to be thinking about this in terms of the number of observations in our data set.

13:04 So I'm going to be focusing on this ratio which we call C for complexity which is just the number of parameters relative to the number of training observations that I have at my disposal. And traditional statistics such as that box Jenkins idea that I gave you at the beginning, the tuki principle of parsimony is based on behaviors like those that I've shown on this plot here which is as I increase the number of parameters of my model from a really small number to a number that starts to approach C equals 1 i.e. the number of parameters starts to approach the number of observations that I have to train on the performance of a model in terms of outof sample behavior. Right? So out of sample R squar can start to look disastrous, right? Bigger models start to do worse and worse out of sample as we approach um C equals 1 from below.

13:54 What's happening here is that as C approaches one, I'm fitting the training data better and better. when C equals 1. In fact, I exactly fit the training data. When I have the same number of parameters as observations, I exactly fit the training data. When I do that, my beta estimates have an enormous variance. Because of that enormous variance, when I bring this model out of sample, the forecast variance is so large that the R squar becomes extremely large and negative. All right? Now this is per perfectly in line with the the box Jenkins intuition and it's perfectly in line with a lot of the ideas of statistics that we were trained on. But what's less obvious is why this happens when we take C above one. When we start to use more parameters than observations, some very interesting things happen. First of all, the variance of my estimates doesn't continue to rise. It actually starts to recover.

14:56 and the R squar doesn't keep getting more and more negative. It actually starts to recover and eventually becomes positive. Why does this happen? The reason why this happens is not something that we can we can get into in detail um in the length of this presentation, but it's because when we choose a large model, we are performing a particular set of operations that allow us to improve the approximation capacity of the models that we use, but at the same time improve that approximation with a relatively low variance. The low variance comes from a form of implicit regularization that happens inside these models. This implicit regularization is critical to the behavior of complex models and it is in fact the source of the virtue of complexity. All right.

15:49 Now, in order to understand the idea of implicit regularization in detail, I'm going to need to refer you to the original paper that we wrote on this part this topic, virtue of complexity and return prediction that we published in Journal of Finance a couple of years ago. But taking that idea as given, what this ultimately implies for the behavior of financial models is that as we increase the number of parameters in our model, the performance of those models increase in terms of their expected returns.

16:24 Their volatility increases at first but eventually recovers due to implicit regularization. And then when I look at the overall usefulness of these models in terms of sharp ratio, I'm really just thinking about the ratio of these two effects. The sample sharp ratio is improving as I add parameters. Why is this? It's because as I add parameters, the approximation of my model is getting better and better. It's getting to look more and more like the unknown true model. But the variance of the model is not exploding. And that's really the surprising effect, right?

17:02 So the surprising effect is that as variance, sorry, as the number of parameters grows, approximation capacity grows. We sort of knew that would happen, but what we didn't expect was would happen was that the variance would also sort of start to be controlled in this particular region of C greater than one. All right. So the surprising benefit of these complex models is actually coming through this relatively low variance complex models. Again that's thanks to implicit shrinkage. All right. So the way you should think about this is that the virtue of complexity arises from a trade-off and it's coming through the fact that approximation benefits of complexity dominate the statistical cost of having to estimate all those additional parameters.

17:48 All right. So for those of you again that are interested in more detailed intuition on this, I'll point you to the JF paper or to a new paper that Seamian and I wrote recently called understanding the virtue of complexity. All right. What I want to do now is move on to the cross-sectional setting. I want to think about asset pricing models for many assets. So if the if you recall the original formulation I gave you, it was about thinking in terms of predictability for a single asset in the time series. I want to move now to thinking about the joint pricing of many assets in the cross-section right on the right.

18:23 >> Brian, if I if I may just could you just uh explain us uh is there any specific assumption that you have to make on the signals for instance on the coarian structure of the signals that you use on the weakness or or or strong uh uh with respect to this general setting. >> Sure. This is going to take us a little bit of while to get through, but let's talk about it. So the key assumption that I need here for this result is an assumption on the nature of the signals which I'm going to describe in terms of the beta on those signals. Okay, so the assumption here is that these beta coefficients are x anti- identically distributed. So what does that mean? What that means is that before I go to train a model, when I build these nonlinear features in my neural network, I can't tell you which of these nonlinear transformations is going to be more or less useful for the ultimate prediction than another. That's the assumption that I need. Now, what's interesting is that that assumption is more or less guaranteed by the structure of many of the neural networks that we look at. If you think about it, what happens inside your typical vanilla neural network is you take these original predictor variables and you start to distribute their information in an Xanti identical way across all of the nonlinear terms. I'm not sure if you guys can see this thing that's cropping up here, but um so you start to distribute this information in an exantiidentical way across all of these hidden neurons. Now, when you do that, you're in essence precluding yourself from being able to understand whether it's this first nonlinear transformation or the last nonlinear transformation that's more or less likely to be useful.

20:17 So if X anti all of these predictors are identical in their additivity then of course expost will be very different in their additivity. Some will be more useful than others always expost but Xanti that's the assumption that we need. We like this assumption because it is very closely tied to the structure of the types of neural networks that we use in practice. All right. So what that gives you is in essence this idea that as you increase the complexity of your model, the incremental benefit of adding each predictor doesn't start to decay.

20:50 Right? Now of course this curve is going to eventually level off because the information is being right kind of subsumed by many predictors as I add the each each subsequent predictor. But the individual contribution of any one is xanti identical. Right? So that's where that's where um a lot of this result is coming from. Okay. So I want to move on now to talking about the cross-sectional setting. And to think about that the important place to start is to think about the AP. So what I want to do is to recognize that when we write down asset pricing models and we write down small parsimmonious specifications of asset pricing models, we are oftentimes working in essence off of Ross' conjecture. And that conjecture is that there's a small number of factors that govern the joint variation of returns.

21:41 Okay? But I want us to appreciate that this is a conjecture, right? We can talk about the empirical validity of this conjecture. That's going to be important content of this paper, but it is a conjecture. And what we have from the AP is that once we have this conjecture plus some no arbitrage arguments, then we get some pricing relationships that emerge. And we like that model because it's tractable to work with mostly due to its parsimmonious assumptions. But we could start from a different conjecture and that's exactly what we want to do in this paper. So in the AP or a AIP setting, what we're talking about is an AIP as beginning from essentially the other end of the spectrum that Ross started from. So we want to consider asset pricing models with potentially an exorbitant number of factors and each of these factors just like the nonlinear terms in a neural network each of these factors have some ex anti viability of being useful. All right. So this conjecture is rooted in the theory of ML and AI, the idea that complex models tend to outperform simpler models. And so we call it AIP um as a juxiposition with AIP, artificial intelligence pricing theory. But the premise here is that if we think about the history of sort of empirical asset pricing, we can encapsulate it in the SDF framework.

23:04 Right? So if I have some SDF, let's call it M, that M prices risky assets R I with zero pricing error. So this is just a statement of the oiler equation in any SDFbased model, right? And so I can represent this SDF in tradable terms as a portfolio of the risky assets in the economy. Okay? So wellestablished theoretical results in asset pricing tell us that this representation holds um as long as there's no arbitrage in the economy. All right. The big question then is what is this tradable portfolio?

23:42 How does it look? In other words, how do I weight each of the risky assets? What does this waiting function look like? All right. So when we typically build asset pricing models, think about the FMA French model as a starting place. We're thinking about this as a conditional waiting function. The weight we put on say Microsoft or Caterpillar Tractor or whatever asset we're considering is going to be a function of the observables for that asset. And we can represent the FMA French SDF in a very simple sparsely parameterized model a parsimmonious model. Right? What the fa model in essence implies is that the weight on each individual asset is dictated by the size of the asset and the value of the asset. Or if you look at the more recent incarnations of the model, it's going to look like it depends on a couple more characteristics.

24:35 So what does this accomplish for us? It gives us an SDF that's extremely parsimoniously parameterized. And of course that has some attractive features because parsimmonious models are going to have coefficients that we can estimate with low variance. But it also has some detracting features. And the detracting features is that this is just a human-chosen model specification. And the correct specification for this SDF waiting function might look very very different. And so what we want to consider in this paper is just a different specification for that waiting function. Rather than restricting W to be parsimmonious, we want to blow W up.

25:16 We want to include an enormous amount of conditioning information in there. And the way that we're going to do it is that we're going to make W a function of a large conditioning set in a very nonlinear way. And as you might guess, we're going to do this with a neural network. All right. So the formulation that we have in this particular model is in essence the analog of the formulation that I was talking you through in the earlier slides is I'm going to begin with conditioning variables. Think of these as stock level characteristics.

25:46 They might be size, value, momentum and then I'm going to produce a bunch of new nonlinear transformations of those characteristics that I call S. All right. So think about this first first term S as taking all of the characteristics for a stock wrapping it through some new nonlinear tweak and just producing a new observable function of these characteristics that I have at my disposal. There are of course infinitely many such transformations I I can make when I build these nonlinear functions. It's as though I've just taken one data set of stock characteristics and converted it into another data set of stock characteristics using some functional transformations.

26:32 Now when I do this, what I've done is in essence taken the original characteristics and tried to capture a whole lot of their nonlinear attributes, their nonlinear behavior. That's what each of these numerous terms in the hidden layer of a neural network are representing. Now once I do that and I rewrite my SDF in terms of this approximating function for W, what I see very quickly is that I have all of these characteristics that I've built. They may be very highly numbered. They might be thousands or maybe even millions of nonlinear characteristics.

27:07 I plug in my W function and now I have these characteristics interacted with the set of risky asset returns R. And of course this is no surprise. We see that this S primer R is a collection of what we typically refer to in the asset pricing f asset pricing literature as factors right these are just dynamic trading strategies they are characteristic managed portfolios all right so what we can do then is say that any highdimensional SDF is going to have a representation as just a combination of a highdimensional set of factor returns right this is a very close analog to the FMA French model. The FMA French SDF is just a linear combination of a handful of characteristic managed portfolios. So all we've done is ex expand out the specification of our SDFs so that we can include much more complicated versions of the usual conditioning characteristics than we typically use.

28:08 All right. So in the interest of time I want to show you a couple of results. First just to give you a little bit of a clearer insight into the way that this model construction is actually processing when I build my factors in say a traditional from a French environment what am I doing I'm using the characteristics as portfolio weights so one of my characteristics might be booktomarket ratio so I take the collection of booktomarket ratios for all stocks I rank and center it and then I use those as portfolio weights and I aggregate up all the risky asset returns according to those portfolio weights.

28:47 Then we call that a value factor. I do the same thing with size characteristic with a profitability characteristic with an investment characteristic and so forth. And then finally what is the SDF in a from a French model? It's just the mean variance efficient combination the tangency portfolio of these four factors five factors. All right. The complex analog in the of this is to say listen instead of just using plain old book to market why don't I produce a new characteristic it can depend on all of the characteristics in very complicated ways and in nonlinear ways I can produce many such characteristics by just varying the ways that I add up all of these usual characteristics passing them through a nonlinear activation function and now I can have perhaps thousands or millions such characteristics. They are not using any different data than the former French model. They're just using the data in a different way, a different nonlinear representation.

29:50 By the time I build out the factors using all of these new characteristics, it just means that my time series of factors is going to constitute about well in this case 10,000 different factors. My SDF then will be a mean variance efficient combination the tangency portfolio of this highdimensional factor set. Now all right now of course because we're using highdimensional models here can no longer use in sample evaluation. In sample valuation is going to be just mechanically scaled up in proportion of the number of parameters that I put in my model. So all of the evaluation that we're going to be looking at here is of course out of sample. But that goes without saying when we're doing um when we're doing empirical asset pricing work with machine learning. All right. So lastly, I'll just show you a couple of of of empirical results and then we'll wrap up. So the the data that I'm going to work with here is really a lot of usual suspect type of data. So I'll look at monthly stock returns in the US starting in 1963. The conditioning set of characteristics is going to include 130 characteristics uh from my paper with Tyson and Lassie Person in 2023.

31:01 And then we're going to look at out of sample performance both in terms of SDF sharp ratio and in terms of the SDF's ability to price um a variety of test assets um out of sample. Okay, so here's a very quick glance at the performance of complex SDFs as we go out of sample. So again on the x-axis I have C which is the number of parameters in my model or equivalently the number of factors in my model. So as I increase the number of factors from a small amount think of a two three four factor model these have depending on the amount of regularization that I use somewhere on the order of one to one and a half as the mean variance efficient portfolio sharp ratio. This is typically what we see in the asset pricing literature. As I add more and more nonlinear transformations and there therefore more and more nonlinear factors, what I end up with is a big increase in out of sample performance. Note that as I move from small C to higher and higher C, higher and higher complexity, I'm not changing the information set. I'm just changing the model specification to include more and more parameters. All right, so sharp ratios are increasing in parameterization. Pricing errors are decreasing in parameterization.

32:24 All right, let's do a little bit of model comparison. So what I just showed you was the behavior of a complex stoastic discount factor. This is the sharp ratio of the complex stoastic discount factor out of sample. I want to compare that with the out of sample sharp ratios of some of the standard parsimonious factor pricing models out there. So we have the FMA French five factor model plus momentum. It gives you about a.7 sharp out of sample. We have the standby yuan mispricing model. We have the hosu is q theory based model.

32:56 We have Daniel at all. We have burila shanken. All of them are giving you something on the order of 75 to 1.25 in terms of out of sample sharp ratio. So when I say that machine learning in finance does not just improve performance in percentage terms, it's actually improving performance at a much more dramatic scale. This is what I want to emphasize. These are the types of differences that you can get again just from introducing nonlinearities into your specification.

33:26 In panel C, I'm showing you pricing errors. So um I'm looking at a couple of different uh models. Those are specified by the different bars and then the different groups of bars are different test asset sets. So what if I use the original JKP factors as test assets? What if I use an enormous number of test assets that correspond to some complex representation or even just the original FMA French 25 size and value portfolios? What you see is that the pricing errors coming from the complex model in all of these situations are much smaller than you get from parsimmonious models. Right? So this is a nice nice kind of comparison.

34:05 The thma French specification was originally designed to answer this question about how do you price these complicated portfolios of 25 size and value sorted assets. But even with a complex model we can we can uh beat this parsimmonious model that was sort of designed for this setting. The last thing that I want to do this will be my very last side slide. So, I appreciate you appreciate you bearing with me. In everything that I've shown you so far in these empirical results, I've used these 130 stock characteristics from the JKP paper. Um, what I want to show you now is the behavior of complex models if I literally restrict the inputs of the model to just the four conditioning variables in the FMA French model. All right. So, this is using I'm sorry, five conditioning variables. size, value, profitability, investment, and momentum.

34:59 So, what I'm going to do is I'm going to look at the original farmer French model. Then, I'm going to compare that model's SDF performance to a model that uses the exact same set of characteristics, but represents that in tens of thousands of nonlinear transformations of just these five characteristics. And what you see is again this is your your 7 sharp ratio for the FMA French model plus momentum. What happens if I just take FMA French data and convert it into a complex specification? You see that the sharp ratio ranges from something like 2 to 2.5 depending on the amount of shrinkage that I use in that model.

35:39 Okay. So I have there's more in the paper thinking about the economic underpinnings of what's going on in this SDF. We don't have a lot of time to go into it here. Um so I'll just wrap up with a couple of brief conclusions. So the first conclusion is that asset pricing as a research field, asset management as a practitioner field are in the midst of an ML boom. We are in desperate need of theoretical insights into the properties of ML models for these types of applications and that's what we aim to deliver uh in the sequence of papers. And so what we show, this is both theoretically and empirically, is somewhat contrary to conventional wisdom perhaps to say the least, that higher complexity improves model performance um pretty much across the board in terms of the analyses that we looked at. And the best way that you can think about this is that the benefits of improving the specification of models by adding parameters outweigh the costs of statistical variance, sampling variance that comes from having to train those additional parameters. An important caveat is that this is not a license to add arbitrary predictors to a model. Think of it more like the FMA French setting that we just went through. Start with what you believe are plausibly useful predictor variables.

36:56 conditioning variables. Once you have an economically plausible set of conditioners, then we recommend using those conditioners in a highdimensional rich nonlinear way. That's where the benefits of of of complexity come from. All right, so I'll stop there. Thanks very much for giving me an extra minute or two. >> Thank you very much, Ryan. Uh, great talk. So now I pass over to my colleague, Professor Ramanipal. >> Thank you. Thanks, Brian. That was a excellent talk. You've also broken the record for the maximum number of participants at one of these talks. We have more than 300 people listening in.

37:37 Uh I think it's very clear that the work you present is a great departure from the way uh at least in finance we've been thinking about for the last 70 years. And so uh there's a lot of soulsearching amongst uh people who believe in complexity and people who uh find it difficult to leave the world with which they grew up uh believing that uh simple models must beat complex models. So uh Brian my first question to you is about the definition of a factor. So like if you look at Chamberlain and Rothschild they define a factor as something that explains the co-variation of asset returns you know so something with a large value and so when I look at the kind of data you've looked at I've rarely found uh more than three or four factors so you know even before you start doing any asset pricing if you just do purely statistical analysis is you find that three or four factors explain 99% of the variation. So all the other factors that you are building what exactly do they explain?

39:04 >> Right? This is an excellent question. So just in terms of of nomenclature right when I talk about a factor a factor is a characteristic managed portfolio in our environment. you raising the question about what is in essence an igen value or an igen vector in a covariance matrix of assets. We spent an enormous amount of time talking about this in the paper. So there's a very close mapping between what's happening in complex models and what's happening in this this space.

39:35 Right? So when we talk about factors and we talk about complexity in terms of the number of factors, we're talking about just building many characteristics upon which we can form sorted portfolios. Now the joint behavior of those sorted portfolios is going to be critical for understanding the behavior of a complex SDF because that's going to really rely on the covariant structure of these Fs. So what you get out of the theory that we build in our paper is really a theory of understanding the co-variance structure the en value and igen vector structure of a covariance in a complex environment. One thing that's really important and you touched on this point is the question of inference about the number of true factors in a coariance matrix and this comes back to right my my original motivational point um versus a Apt is saying well we have a coariance matrix it has a couple of dominant igen values and the rest are are much smaller whereas the AIP conjecture can be viewed as actually the distribution of igen values is much more even right there are many many more factors that can be thought of as having an important igen value associated with them which is true well you might say well why don't we just look at the empirical coariance matrix of factors well You can do that if you are in a parsimmonious environment. If complexity is low, if complexity is high, it actually begins to be very very difficult to infer the true number of factors because you don't have enough data points to estimate the igen value distribution accurately. In fact, random matrix theory tells us that these igen values in a complex environment get to sort of get all smeared out. it's very difficult to infer anything about their structure. We have a big discussion about this in their paper. So I think that what we have in the literature is a lot of assumptions on the econometric environment that allows us to measure the true number of dominant values. But we believe that in the environment that we're studying actually those assumptions are violated. Right? So I would just kind of conclude with a caution about us interpreting the prior literature correctly as implying that there is clear evidence that there's a small number of common factors in the cross-section of returns. I think the evidence is much less clear about that.

42:09 >> So like uh I'll say you don't need any assumptions actually because you can just plot the EN values and just do a simulation where there are no systematic risk factors. compare the two plots and you can see that the number of latent factors is actually quite small but >> absolutely we we have this analysis in the paper. This the exactly the right way to look at it. Simulate a lowdimensional factor structure and then calculate the empirical coariance.

42:39 >> Yeah, >> the empirical coariance structure the value distribution is going to look very very different than the true distribution. And we have a nice plot in the paper that shows whether you simulate from a small number of true factors or a very large number of true factors in a complex environment the igen value distributions are right on top of each other and that's exactly as the Marco Permer uh theorem would predict. So this idea that you can infer by looking at the empirical the empirical igen value distribution is just not true. we that's a that's a statement that's true if you have um a a much more parsimmonious environment than the types of environments that we're that we're considering here.

43:22 >> So Brian, my second question to you is uh taking advantage of your uh experience at AQR. So in terms of trading strategies uh what do the complex models generate like in terms of trading volume price impact short positions leverage how easy is it to implement these policies? >> Sure. Um it's an interesting question. So let me just say that my role in the asset management industry is to oversee a team of researchers that build machine learning based portfolios. Um I think the existence of this role probably speaks to the effectiveness of these types of approaches. um at the types of frequencies that we often look at in the asset pricing literature, things like monthly rebalancing um and with the types of transaction costs that we sometimes consider in the asset pricing literature, uh you can expect these types of analyses um to deliver excess performance above simple models in a net sense.

44:32 So what we see in the literature is not um so what we see in the literature in gross terms is of course an overstatement of what's possible because what's happening when we when we trade in the academic setting is we're trading an enormous universe of assets that includes small caps and micro caps. Many of those may not be tradable. Many of them are certainly not shortable. Right? And this is going to be a fairly short intensive strategy. But if we were to sort of reduce this to a set of assets that are highly tradable, in fact, I think I might have um Yeah. So here is a a breakdown into different size groups.

45:15 So here I'm looking at SDFs. If I just look at about the top 500 stocks every period, certainly a highly liquid universe and everything in there is shortable. you can produce something that looks like a sharp ratio on the order of 1.5ish let's say right I think this is probably closer to something that that that would be reasonably imple implementable but the key point is not what is the level of the sharp ratio the point is that as I look across complexity all of these models have access to the same assets right they have access to the same information they have access to the same trading costs right but it's the high complexity models that are more effective across the board.

45:56 >> Thank you, Brian. Like maybe we'll give Thomas a chance and then there are some questions uh in the chat. >> Yeah, perfect. Thank you, >> Thomas. >> Thank you very much. Yeah, thank you very much for this very insightful presentation on a very topical subject. I think it's very interesting for us as a students to to to get exposed to this kind of papers. So, thank you for that. The first question uh on my side would be well you suggested that we should use the most complex model we can compute.

46:35 How does this principle apply when for example we believe the underlying data generating process is very simple for instance over a very short period of time and in practice uh how do you see the trade-off between the potential currency gains from complexity uh versus the costs which can be you know computational or otherwise uh of implementing such uh models.

47:07 >> Yeah, great question. Okay, so um first of all, I would say that it's it's an assumption that the underlying data generating process is simpler at high frequencies. I don't I've never seen any evidence of that fact. Um so let's leave that part aside and let's think about what are some of the costs that we face of of implementing highdimensional models. Well, one is of course compute cost. I think that's why I was sort of talking about you should compute a large model if you can compute it right in reality compute is not free and so in practice you can expect that asset managers are not building infinite dimensional models or nearly infinite dimensional models right um there's going to be some interior solution that's that's optimal um because of compute considerations the other thing that might be a little bit more subtle are the operational complexities of using complex models Okay. So, one such operational complexity is that you need to be able to um develop code bases that can be used throughout an organization. You have to be able to diagnose issues in the model sort of in real time. Um and then if you think about kind of interaction with investors, if you're as a if you're an asset manager, you have a client base that you need to communicate with. And so questions about communicating what your model is doing rel uh to your investors is an important consideration.

48:33 Now I think all of those are are addressable um but they do on the margin point towards smaller models right they're going to constrain ultimately the size of your model but where you land in that costbenefit trade-off. The benefit of using a big model is um substantive enough that the costs of using high complexity models um are not going to are not going to um sort of offset the benefits until you're you're quite deep into the high dimensional spectrum, right? Um, so you can think about asset managers that are using machine learning models as very likely having models that have at least thousands of parameters if not many many more. Um, but perhaps you know it's we don't need to worry about them yet using models you know on the order of GPT.

49:29 >> Okay, thank you. And to just jump uh on on your answer, um did you ever apply this kind of uh very highly parameterized model in your uh in your position for example at AQR? And when for example you had uh performance attribution to to to tell to clients, how did you manage it? >> Yeah. Right. Um it's a it's these are good questions. You know there there are some compliance issues with with giving answers. I gave a little bit of a suggestive suggestive answer to to Ramen's question on this topic which is um we use these models. We like these models. Um they are important contributors to the behavior of our trading strategies. Um but I can't give any kind of breakdown in terms of what their what their performance looks like on the practical side.

50:31 Okay, thank you. Um, I >> do do you have another question or we we move to to the Q&A? >> I can do the the floor to the chat. I don't see if there >> Okay, so we have we have several questions from from the chat box. Um, there is one uh I pick Haman if you want to pick another one after me. Um, up to you. Uh the first question doesn't your C uh your complexity ratio P under T imply that we can simply increase complexity by just decreasing the number of observations. Uh then you are implying that looking at just the latest available data would be better than looking at longer time series. So so just interpretation of your complexity ratio which is a ratio of the number of parameters with respect to the number of observations. Sure. Right. So, um, yeah.

51:27 So, I encourage you to take a look at the paper. Um, what I'm showing you here are comparative statics that can be viewed as the role of changing the number of parameters holding the the number of data points fixed. Okay. So, these comparative statics are very useful for that type of analysis. The theory is much more complete than that. In the theory, if you start to increase complexity by lowering the number of observations, now you have two offsetting effects, right? There are some benefits. Let's say that you are um you're starting from a point where t equals p. Notice that there's a lot of destructive behavior from t equals p.

52:08 That's sort of the worst case scenario when it comes to model complexity. So even if you held P fixed and shrunk T in this very narrow region, you might actually start to improve model performance. But that's not really the point. The point is that as you shrink t, what's going to happen more generally is that all of these complexity curves, they're going to start to bend down. They're going to start to flatten out. Right? Information makes models better.

52:37 And that's almost always true. The only time that's not true is when you're in this sort of this um domain of attraction around C equals 1. Anytime you're out here, adding information by adding more observations is going to improve performance. Right? So the best way to think about this is that these curves give you comparative statics or the role of P holding T fixed. If you wanted to add the role of t, you'd have vertical shifts in these curves. More t is going to shift these curves up. Less t is going to shift these curve curves down.

53:18 >> So, >> let me ask a follow-up question to what Thomas had asked. Uh, a a lot of these models are like black boxes. So, how do you deal with regulators and clients to communicate to them what you're actually doing? Yeah, ab this is a great question. So, you know, when I think about the sort of the ecosystem of the asset management industry, in particular, the quant asset management industry, um, attribution, right? So, let's call this broadly attribution, communicating to people what your model's doing. Um, attribution is an important topic. Um but what I find is that attribution is actually dictated by the investor's demand on how information is presented to them because these are we can call them black boxes but they're perfectly defined. They're perfectly specified. They're completely specified. There's no question that I cannot ask of my model. anything about my model I can inspect to an arbitrary degree of precision right there's nothing stochastic about it or anything like that so if if the investor wants to know okay how much does this model expose me to inflation risk well that's a very clear attribution that I can do because what I can do is take the model output and projected onto inflation or what does my model do um when there is some sort of meme media Right. Well, if meme stocks are going crazy, I can look at how my model behaves relative to that set of stocks by looking at the correspondence and my model positioning versus those stocks.

55:02 Again, it always boils down to some form of projection of a complicated model onto a simpler model. So, the way I think about it is that whenever you're doing asset management with machine learning, you really sort of have three models floating around. You have your full-blown machine learning model. You have your simpler model that people like to look at and interpret. And then you have a model mapping those two together essentially accomplishing the the projection of one onto the other.

55:32 >> Thank you, Brian. There's another >> interesting question. uh according to the universal approximation theorem uh twolayer neural networks create all possible combinations of nonlinear features. So do we need to look at anything else or is that enough then? >> Yeah. So this is obviously a very interesting theoretical point. It's a point that's been studied a lot in the machine learning literature. So of course the you know some of the earliest forms of the universal approximation theorem as they applied to neural networks came out in the 80s and they were focused on on shallow one- layer neural nets. Since then we've seen the beneficial performance of narrower and deeper networks right and there's been some theoretical advancement in trying to understand the approximation benefits of depth. Um, so one thing that I do not want you to take away from anything that I'm doing here is that I advocate using a shallow wide neural network. That is not the case. That is the tool that I used to demonstrate the benefits of complexity in the examples that I showed you. What I do believe is that high complexity models, generically speaking, are beneficial.

56:51 I'm not making a specific model recommendation. In fact in practice I very rarely use shallow networks. >> There is one question with respect to to to the virtue of complexity everywhere. So are there asset classes or markets where the virtue of complexity is weaker? So the question is it asset dependent or is it something that apply on any kind of asset classes? >> Yeah. So um I've not seen any qualitative difference in behaviors like this across asset classes across different types of assets, industry size, volatility, right? And you can sort of see this here as we look across the size spectrum. The qualitative patterns are essentially indistinguishable. If I if I, you know, sort of blocked out the the tick marks on the vertical axis, you'd have a hard time knowing which universe we were looking at. Um, so this is a I think of this as a fairly universal pattern. Now, in the paper, understanding the virtue of complexity, we talk about conditions under which these curves can can change shape. So if you're interested in details about what actually gives rise to the specific shape of the curves that we see um I encourage you to take a look at that discussion. I think um there's a pretty uh beautiful way to characterize the origins of these curves in terms of two attributes um yeah which we call at uh which we call um concentration and alignment. But again that's that's too much of a conversation for us to have in a Q&A.

58:31 Thank you, Raman. Do you want to pick a last question or are you done? >> We all all set. Thank you. >> Okay. So, thank you very much. So, thanks uh thanks a lot. Thanks Brian for this talk. Thanks Raman Thomas. This was a great session effectively. Uh we had a lot of attendees and also a lot of questions more than 15. Uh just before we wrap up uh we will be back next week with one coer of Brian. So last Peterson will be there to talk about the ginium.

59:02 So this will be another topic and we'll examine the true green in different asset classes at different time. So thank you very much and see you soon. Thank you Ryan. This was great. >> Thank you so much. It was my pleasure. Take care. Thank you all. Merciu. Bye. [Music]

Summary

Professor Brian Kelly discusses the implications of complexity in asset pricing models, emphasizing the advantages of highly parameterized models over simpler ones. He argues that while traditional statistical teachings favor parsimony, modern machine learning applications demonstrate that more complex models can yield better out-of-sample performance due to implicit regularization.

- Complexity in financial modeling challenges traditional views favoring simpler models.
- Highly parameterized models, like neural networks, can outperform simpler models in return prediction.
- The principle of parsimony often leads to skepticism about complex models, but empirical evidence shows significant performance gains.
- Theoretical insights suggest that increasing model complexity generally improves performance metrics like Sharpe ratios.
- Implicit regularization in complex models helps manage variance, countering concerns about overfitting.
- The discussion extends to cross-sectional asset pricing, advocating for models that incorporate a broader set of factors.
- Practical applications in asset management reveal that complex models can effectively enhance trading strategies, despite operational challenges.
- The findings suggest a need for a paradigm shift in finance towards embracing complexity in modeling approaches.
© transcribe · For agents Built with care and craft by Gokul Rajaram