NVIDIA Corporation (NVDA) Earnings Call Transcript & Summary
February 9, 2023
Earnings Call Speaker Segments
Markus Pelger
attendeeThank you very much for listening to this presentation. This paper, Missing Financial Data, is joint work with Svetlana Bryzgalova from LBS; Martin Lettau from UC Berkeley; and my Ph.D. student, Sven Lerner, from Stanford University. So this paper is about firm fundamentals or firm characteristics. These are crucial for investment and asset pricing. We use this firm fundamentals to build investment strategies. For example, by sorting stocks based on these characteristics or by predicting returns given this whole universe of firm characteristics. We use this firm fundamental to build asset pricing models, for example, factor models. And we use firm characteristics to construct test assets for our models. This could be double-sorted portfolios. Now everyone who has worked with firm characteristics know there's 1 fundamental problem. Not every company has all firm characteristics observed. In fact, there's a lot of missing data. Now the key questions that we addressed in this paper is, does missing data matter? And if yes, how should we deal with it? Now the first answer is yes, it matters a lot, otherwise, I would not give this presentation, is the way the literature has dealt with missing data is very problematic. So 1 standard approach of how papers and people in the industry have dealt with missing data is only to use a subset of firms that have fully observed characteristics. Now we will show that this is going to lead to a massive sample selection bias because companies that have observed firm fundamentals are different from companies that have missing data. A second approach is to use some form of ad hoc imputation. For example, a cross-sectional average or some observed past values. Now we are going to show this is also going to lead to biased imputed values so that it will strongly negatively affect any follow-up or application. Now what we do in this paper is, number one, we shall key facts about missing characteristics; number two, we provide a solution, namely, a way to impute missing values; and number three, we show what the application now for asset pricing and investment. I want to highlight what we study as a broader impact. So number one, we need this firm characteristics not only for asset pricing or investment, but there are other applications, for example, in corporate finance, where our results will be useful. And number two, the message and insights that I present here, while we use firm characteristics as our main object for the study, the same logic applies also to other types of data. For example, for ESG data, we have a lot of missing values, the same for international data. And what we present today here can be transferred to that type of data as well. Now let me start by establishing 4 stylized facts on missing firm fundamentals. So number one, missing data is prevalent. That means most companies and most characteristics have missing observations. This is going to affect small and large, young and mature and profitable and distressed companies. Number two, and that is what the big elephant is in the room. If you require a firm to have multiple characteristics observed at the same time, a lot of data missing. In fact, if I would only look at a subset of companies that have all my characteristics observed, I would throw away over 70% of the data. So this would represent over 50% of the market capitalization. This is huge. Now please keep in mind, there are a lot of applications where we require a company to have all characteristics to be observed. For example, if you run some form of regression of returns on characteristics, we need a fully observed backward characteristics. The same for a lot of machine learning application that try to predict returns. Number three, data is not missing completely at random. So companies that have missing data are systematically different from companies that have observed data. So we observed systematic patterns in missingness so they will be clustered over time. So if something is not observed today, it's more likely to be also missing in the future. And there will be clusters and cross-section of characteristics. It means certain characteristics will be more likely to be missing together. Importantly, companies that have missing data are more likely to have more extreme realizations missing and their characteristics are different. For example, they are more likely to be small companies. Number four, returns depend on missingness. In fact, investment strategies that use all stocks with imputed values will lead to more profitable out-of-sample results. Then if you would restrict these investments -- restrict these investment strategies to the subset of companies that have everything observed. We established that there are 2 key elements of -- or effects of how we deal with missing data. So number one, there's a selection bias if we only use a subset of companies that have everything observed. And we show that this has an impact on prices, even for simple anomaly strategies or simple long-short factors. Number two, there's an imputation bias if you use simple ad-hoc methods like the median to impute missing values. And that will strongly bias risk premia estimation and follow-up applications. So overall, there are widespread implication of how we deal with missing data in the context of asset pricing. So the second part of the paper is of how we should impute missing values. Now imputing missing data is challenging because there are 2 steps that we need to solve. Well, number one, any model for imputing missing data starts with a good model for characteristics itself. And here, we want to make sure that we are not missing a relevant information. We want to avoid an omitted variable bias for the model that we built for characteristics. Number two, we need to estimate this model on the partially observed data, but then this model has to be valid on the missing data. And now keep in mind, again, data is not missing completely at random. There's systematic patterns of how data is missing, and that needs to be taken in account to make the imputed values valid. So what we propose is a cross-sectional and time-series factor model for imputing missing values. So there's 2 components. There's a contemporaneous cross-sectional latent factor model. This will, for example, capture relationships like that small stocks are more likely to be valued stocks. So it will capture the dependencies in contemporaneous characteristic observations, and it's going to do it in a data-driven way. Number two, there's a lot of persistence in a lot of characteristics, and we will capture that with a time-series factor model. So by combining these 2, we leverage this 2-dimensional information. What is really important is that our approach allows for very general endogenous missing patterns. So it will be valid if missingness depends on time, on stocks, the characteristics or even the factor model itself. So the main -- it's really important that you understand this. Our model will be valid even when missingness satisfies all these complicated missing patterns that you observed empirically. And it really separates our approach from other approaches in the literature. So overall, we have a data-driven, transparent and very simple-to-implement approach that works extremely well empirically. So what we've shown, a comprehensive empirical comparison study that our imputed values are 40% to 50% better in terms of imputation error out-of-sample than existing benchmarks. We also show that it is important to include both this time-series information for more persistent characteristics and this contemporaneous cross-sectional information to get reliable imputed values. Now we will -- for the data set online that will be available for researchers and practitioners as well as the code so that you can use our imputed values in your own work. Now let me start by documenting stylized facts about missing data. So the data set that we will consider here is very standard. It's using the standard CRSP/Compustat universe. We have over 50 years of monthly data. We consider 45 characteristics based on value, investment, profitability, intangibles, past returns and trading friction information. And we will normalize our characteristics as centered rank quantiles that is the result of a lot of generality, as we discussed in the paper. And we also show how to transform imputed values back into raw values. One thing I just want to mention is some characteristics are updated monthly, as are characteristics that are based on accounting variables might be only updated quarterly. Now with our imputed values, we will also get monthly time-series for quarterly updated characteristics. But when I present statistics about the degree of missingness, I will take into account the updated frequency. So only tell you how much quarterly values are missing, for example. So what I really want you to take away is that I'm using a very standard data set, the type of data set that has been used extensively in asset pricing and investment studies. Now let me show you how much data is missing. Here, the black line shows you the number of stocks that we have in our data set. So at its peak, we have around 8,000 stocks. And this varies over time. At the end, we have around 5,000 stocks. Now if you look at the blue line, that shows you how many of those stocks have, for example, operating profitability observed. At the beginning of this sample, there was more missing data. But at the end of the sample, around 10% to 15% of the stocks do not have operating profitability observed. The light blue line shows the missingness for property over asset ratios. Here, you can see that throughout our sample, over 50% of the companies do not have these characteristic observed. The main takeaway here is that the degree of missingness vary substantially over time and for different characteristics. So there's a lot of variation. Now the big elephant in the room is if you require multiple characteristics to be observed at the same time. So what I'm showing here with the blue line is the number of companies I would keep in my sample. If I would only keep those that have all my 45 characteristics observed at a specific point in time, I would throw away over 70% of the data. And that corresponds to over 50% of the market cap. This is huge. Now this is not driven by 1 specific characteristic. If I allow up to 3 characteristics to be missing, and that would be the black line, and these characteristics can be different for different stocks and different for different time periods, I would still throw away around 50% of the data. I would need to allow up to 35 characteristics to be missing, that would be the top line, to end up with a data set that is around 90%, 70% coverage. Now again, keep in mind, there are a lot of applications where you require a lot of these characteristics to be available. These are, for example, regressions, cross-sectional regression characteristics, most machine learning applications to predict returns or build asset pricing models, et cetera. Now I want to talk more about when are characteristics missing. So here, I show you all my 45 characteristics and the percentage of missing values. You can see that for some, we have 50% missing values. For the bulk, it's around 20% and up to 10% for a certain class of characteristics. What I want to highlight here is that missingness can happen at the beginning, in the middle or at the end. Now some of these missingness can be mechanical. For example, if we have a new company, a young company entering our sample, it does not have prior data. So characteristics that are based on the history of observed values might not be available. Similarly, companies that are at the end of their life might also have mechanical missing values. And the reason I want to highlight is because it will affect the way, the type of information that we can use for imputation. So if a company has missingness at the beginning, we cannot use prior information for the imputation. And the next point I want to study, which stocks have missing observations? Are they different from those that have observed values? Well, here, I show you the size quintiles for different types of missingness. So the black line is overall missingness. And you can see that the 20% smaller stock, the first quintile, have more missing data in the stocks than the largest quintile. Now if you look at the light blue line, this is for investment, you see that among the 20% smaller stocks, around 40% do not have investment information. But among the large-cap stocks, around 20% have missing values. What you can see here is their complex interaction effects with size. And broadly speaking, companies who have observed values are different from companies that have missing values. Now what are the type of realizations that are missing? Now this is obviously a tricky question because, by definition, I do not observe the values that are missing. Now what I'm showing you here is the average characteristic value that companies have. For example, the light blue line shows that companies that are, on average, in the lowest investment quintile have 50% missing values. Companies that are, on average, in the highest investment quintile, also around 50% missing values. But you have only around 25% missing values if you're on the middle quintile. So these kind of U-shaped patterns call for all characteristics. They might be more or less pronounced. So what this is indicating is that companies that are, on average, in a specific group, more extreme characteristic groups are more likely to have missing values. And that is evidence that the more extreme renovations are more likely to be unobserved. And we will call this endogenous missingness. And I just want to denote, this is a very challenging statistical problem to deal with. Now in order to build a model to impute missing values, we need to leverage dependencies and characteristics. And I want to highlight 2 types of dependencies. So first, a lot of firm characteristics are persistent. What I'm showing you here with yellow -- this orange line is the autocorrelation on a monthly level from my 45 characteristics. What I'm showing you with the blue line is the annual autocorrelation. But you can see that a lot of characteristics are very persistent, have very high autocorrelation values. That means if I observe prior values that is very informative, I want to use that information for imputation. Even if those prior values are far in the past, like a year ago, it might still be quite valuable. Now the second information I want to use is contemporaneous information. So what I'm showing you here are pairwise correlations and characteristics, averaged over time in stocks. One example, what you can see here is that small stocks are more likely to be valued stocks. Now there are 3 takeaways from this graph here. So number one, there's a lot of dependency between the characteristics. You cannot -- that means we want to take advantage of all of this dependency. Number two, below of these clusters, it means that if I observe 1 value within the cluster, I can use it to make very good inference what the other potentially unobserved values should be. Now there are a lot of these clusters. But a priori, I do not know where these clusters are. That's a very -- this is not obvious how the dependency looks like. So ideally, I have some form of data-driven approach that can find all these clusters. And given these statements, a good solution would be a latent factor model. So if you want to model these kind of clusters, you can do it very well with latent factors. So a good model for imputation should ideally take into account this contemporaneous cross-sectional dependency and also the time-series dependency. Now this brings us to our model framework. So the data that we deal with is a so-called 3-dimensional tensor. So we have around 5,000 stocks that are observed for over 600 months for 45 characteristics. And my goal is now to estimate a low-dimensional model that can capture all the dependencies in the 3-dimensional tensor. As I mentioned before, as a baseline model, we use centered rank quintiles. That is the right object to use given that it's more stationary in the time and cross-sectional dimension. But to some degree, without loss of generality, because we can always map it back into the overall characteristics space. Now let's first start by fixing a specific period in time. So then I'm only left with what I call the cross-section and the characteristic dimension. We have a 2-dimensional matrix left. So in specific months, I might have 5,000 stocks for 45 characteristics. Now I want to capture the cross-sectional contemporaneous dependency. The way we are going to do it is with a latent factor model. So we assume that there is a small number of latent factors that captures all these contemporaneous correlation. So we will have characteristic factors F and characteristic loadings Lambda. Now if I do not have missing data, I could estimate this latent factor model with principal component analysis. So essentially, I would calculate a characteristic covariance matrix. And then I use the largest eigenvectors of the -- then use the eigenvectors of the largest eigenvalues to estimate the factors and loadings. Now obviously, I cannot do this in the presence of missing data. So what we are going to use will be the approach that I have developed in my Journal of Econometrics paper with Ruoxuan Xiong. There's a general approach to estimating a latent factor model in the presence of missing data. And I want to emphasize again what makes it very unique. It is going to be valid under very general missing patterns. So here's how it's going to work. So first, we will only use the observed values to estimate this characteristic covariance matrix. So more specifically, if I have 2 stocks, I and J. So then I will see how similar company I is to company J by looking only on the characteristics that are observed for both companies. And then I have to use only those values to estimate a characteristic correlation. Now for different companies, I and J, I will use different sets of characteristics to calculate their similarity. Now this will give me my characteristic covariance matrix. Then I will estimate my characteristic factors as the eigenvectors of the largest eigenvalues. Now given my characteristic factors, I run a regression of these factors on the characteristics to recover my loadings. Now this will be a weighted regression. W is a variable that is 1 characteristic L is observed for stock I at 20. So I'm only using the characteristics that are observed in this regression, but I'm doing some form of weighting in this -- when I run this type of regression. A lot of magic is taking place on how this regression is formulated. Essentially, we'll automatically correct for very general missing patterns. So we show there's some project theory. It means that our estimate will be consistent and also the confidence intervals and as I'm showing normality of this type of estimator in our paper that's published in the Journal of Econometrics. And again, I want to highlight what makes this approach stand out in the literature is that the missingness, modeled as W, can be very general. We allow for heterogeneous missingness, endogenous missingness in the sense that more extreme realizations can be missing. Companies that are small or different in terms of their characteristics can have different type of missingness, et cetera. Now I just want to highlight that a factor model describes this dependency, this cross-sectional dependency very well. What I'm showing you here are the eigenvalues, average over time of this content -- of this cross-sectional characteristic covariance matrix. What you observe here is a type of typical pattern that you have when you have a factor structure. So the first couple of eigenvalues are much larger than the other eigenvalues. Now we discussed in depth in the paper how many factors we need. So what I'm just trying to -- I just want to give you some inclusion here and I show some form of out-of-sample evaluation, so how well is our root-mean-squared imputation error out-of-sample for different number of factors. And the main point is 6 factors around optimal -- for our model, and that is what I'm going to use moving forward. But again, we show in all the details of the paper, and our results are relatively robust. What is nice is that this characteristic factors are interpretable. So we can link those to categories like this value characteristic factor, there's a momentum characteristic factor. And I think, in the results, there's a lot of redundancy in characteristics, and you can describe them with a factor model is actually finding of independent interest. Now we still need to bring the time-series information into our model. So here, I show you how we combine the contemporaneous cross-section information with the time-series information. So given that we have estimated our characteristic factors, I can then include the past observed value of characteristic L for stock I at time T minus 1 in a regression to describe my characteristic model. So this is essentially the same model that I used before, where I got my characteristic loadings, where I only had contemporaneous factors, but now I add the time to this information. Now estimating this model, again, uses this weighted regression that uses only observed values from stacking together characteristic factors and prior value of characteristics. Then I run this regression using only observed values, and this is this weighted type of regression. And the generality of our approach will carry over to this generalization. Now in principle, I cannot only use prior values and combine the contemporaneous information, I could also use future values. So we will have a backward-forward cross-sectional model that uses past, future and contemporaneous information; we will have a backward cross-sectional model that uses past and contemporaneous cross-sectional information; a forward cross-sectional; a pure cross-sectional model that is not using any time-series information; or a pure time-series model that only uses past values. These are all special cases of our model. And of course, then there's a standard approach of just using the previous observed value. This is a special case of a time-series model or using only the cross-sectional median, that's a special case of our cross-sectional model using a zero-factor model. Now we are going to compare all these models to determine what is actually relevant for a good model for characteristics. Now there are different ways of how we can estimate these models. We can estimate them separately for each month. We are going to call this a local model. Or we can stack all our information together in a very large vector and run a pooled regression to estimate our characteristic loadings and also characteristic factors. Now the difference between a global and local model is that a global model should be more efficient because it uses more data. But if you run a local model, we allow for some degree of time variation. But what's even more important, we never use future information. So it will be look-ahead bias [ script ]. Now what I'm going to show you is that our models are extremely stable over time. So using these global models will not be much different from using a local model. So for many applications, for example, if you think about an investment application, you can use this local model to avoid a look-ahead bias. Let me now come to the empirics. So I will show you the quality of imputed values using 2 metrics. One is the root mean squared error. The other one will be a normalized version of the root mean squared error, which is the R-squared. Now everything I'm going to show to you today will be out-of-sample results. So how is that going to work? So we take our data set. We are going to mask values in our data set. Then we estimate our models. We use our model for imputation. Then we compare the imputed value with the masked values, and that gives us some out-of-sample metric. And we will do different types of masking. So empirically, we observe their characteristics, on average, missing in blocks of 1 year. So we use block masking. So we mask 10% of the characteristic values in blocks. We will also have missingness that is completely at random. But again, that is not really capturing what we observe empirically. We will also have out-of-sample logit masking. So here, we will estimate a logistic regression model for missingness that captures all the stylized facts that we have observed and acts as a very accurate description of how data is missing. Then we use this model for masking to create a missing pattern that's as close as possible to what we observe empirically. Of course, we also have in-sample results as well. Now for each model, I will show you the results for a local model that's estimated separately for each month and a global model that pools everything together. And again, the current standard in the literature that the benchmark would be a simple cross-sectional median or just using the previous value. Now here, I'm showing you the main results. These are out-of-sample R-squared values. I'm showing you the results here for block-missing masking by using this logit, this logistic regression model for masking gives you essentially the same results. Now here, you can see that on the -- it's a button, the R-squared values from using the standard approaches. These are top imputation approaches like previous value, cross-sectional median or industry median. All this out-of-sample R-squared are relative to the median. So they tell us how much our imputed value improved versus a very naive imputation. I'm showing you this R-squared for all characteristics combined and separately for quarterly updated and monthly updated characteristics. So the best model that you can use would be a global backward-forward cross-sectional model. And that gives a very impressive performance. It's extremely accurate. Now if you want to avoid a look-ahead bias, the best model that we can use would be the local backward cross-sectional model. It is quite close to the global model that use future information, and it's also extremely accurate. But you can also see that the current standard of imputation like using a median is substantially worse. So in the paper, we have a very extensive evaluation for different types of missingness. So we show how the results look like for missingness at the beginning, the middle or the end for different types of masking, logit masking, missing-at-random, block masking, et cetera. We show the imputation results for the extreme quintiles. So how good are we in including extreme values? We show how it depends on the size of the companies, industries, the results over time, et cetera. So we have a very complete study. The main takeaway is that our model performed really well. Now I just want to illustrate what does it actually mean to impute our values. So here, I'm looking at representative characteristics. This is total assets. That's a very persistent characteristic. And I'm using a representative example stock. In this case, it's Hasbro. And I'm showing you here the actual value for total assets and then the imputed value with different models. And the gray shaded areas are the out-of-sample areas, where I mask a complete block of 1 year. So other characteristics are similar in terms of persistence would be the size of the company, dividend-price ratios, leverage, et cetera. The plots would look very similar. Now there is the observed true value, that would be the orange line. And then I show you the imputed values with our local backward cross-sectional model and our global backward-forward cross-sectional model. So black is our local model and gray would be our global model with future information. What you can see here is that all these lines look essentially the same. And that is because in and out-of-sample, these models all performed extremely well. Now note, if you would use a median for imputation, you would get this green line. And this is obviously completely wrong. It's not only wrong in terms of their relative value cross-sectionally, if you'll create a time-series for characteristics, you use the observed value, then you use the median, then you go back to the observed value, you would completely distort the time series. Now the second example, I look at Tobin's Q. Now this is more volatile than the previous characteristic, but it still has some degree of persistence. Now other examples that would look very similar will be book-to-market ratios, earning-to-price ratios, investment, operating profitability, et cetera. So orange would be the observed values. Blue is a true value when I have the masking out-of-sample area. And then black, respectively, gray are our benchmark model for imputation. Now what you can see is that our models are again quite accurate. And it also captures a type of variation, to some degree, that happens on the out-of-sample period. If you use the median for imputation, again, you would get completely distorted cross-sectional and time-series information. Now what happens when we use a backward cross-sectional model? Essentially, we use the last observed values and anchoring point. And then in between the missing sample, we'll use contemporaneous cross-section information to capture the variation. If you use a backward-forward cross-sectional model, we use the start and the end point for anchoring and then use the contemporaneous cross-section information to capture the variation in between. The third example, I'll show you more volatile characteristics. That's the variance or locally estimated volatility of the stock. Now this is less persistent, right? So most of the information for this type of imputation will come from the contemporaneous information. So pure time-series model would perform extremely badly here. So just to give you an example, if you look at this gray shaded area, you see that the black line, our model actually captures quite well this out-of-sample contemporaneous variation because it has this contemporaneous latent factor model. It learns from other characteristic realizations what this value should look like. And again, using the median as an imputation, would just be wrong. Now I want to also show you how individual characteristics. So here, I have my 45 characteristics. I've ordered them in terms of their persistence. So more persistent on the left and less persistent under the right. And I'll show you, for the different methods what the out-of-sample root means squared imputation error looks like. Now there's a lot of information here. What I want you to take away is that -- so if you use a pure cross-sectional model, that would be the black line, that's here. The best model is a backward-forward cross-sectional model. It's the blue line. And then the local backward cross-sectional model and its variations are here in the middle. What I want you to take away is that, for more persistent characteristics, the gain from using time-series information is the largest. For more volatile characteristics, the contemporaneous latent factor model becomes more important than the time-series part. So that will actually be quite good for imputation. Now to -- and also another takeaway is that the relative ordering of the performance of different models that we have seen for aggregate statistics also hold for individual characteristics. Now here, I want to make more of a statement which information is useful for imputation. I'm showing you here the weight that our model puts on the cross-sectional component and on the time-series component. Now the assorted characteristics such as the most volatile ones are on the left. And the most persistent ones are on the right. So the blue line is my autocorrelation from the different characters. So the orange line tells you how much weight we put on the contemporaneous cross-sectional factor model. And for the very volatile characteristics, we put quite a lot of weight. Now for the more persistent characteristics, we put more weight, not surprisingly, on the time-series information. And so the main takeaways that a good model requires both components, and we get the best results if we combine optimally these information in both components. Let me now come to the asset pricing and investment results. I mentioned before there are 2 fundamental effects. There is a selection bias effect and imputation bias effect. The selection bias effect is how -- what are the asset pricing results if we only take a subset of stocks that have observed values compared to using all stocks but doing a good imputation? So we'll study 3 different types of effect -- we will look at 3 results that shows a selection by the effect. One will be based on portfolios based on the observability of characteristics. We show even univariate portfolio sorts and factors and how the results differ depending on which stocks we include. And we construct a complex asset pricing model, namely with IPCA. What we can show is that out-of-sample investment is substantially better if you include all stocks instead of only using the stocks with fully observed values. And the results will be different in terms of risk premium and Sharpe ratios if we only take the subset of stocks. So it matters which stocks we include. The second effect I want to study imputation bias. It means if I use a suboptimal way to imputing values, like the median, it will distort asset pricing and investment results. And I'm going to show this by estimating the risk premia for different characteristics, and I will explain this in more detail when we come to this part. So let me start with the selection bias effect. So here, I'm showing you the returns on long-only portfolios that either include or exclude particular characteristics. So the black line shows you the average return of a portfolio, long-only portfolio value-weighted, that only takes stocks in the -- have, for example, earnings-to-price ratios observed or book-to-market ratios, et cetera. Now this is essentially a value-weighted market portfolio, and so we observe the average return of a value-weighted market portfolio. Then I look at a long-only portfolio that excludes -- that only have stocks that do not have a particular characteristic observed. What you can see here in the orange line, you get substantially different average returns. Now here, I just look at missingness in the middle. I would get similar results when I look at missingness at the beginning or the end. But then there are also more mechanical effects about missingness, like young companies that have no observed values. The main takeaway here is companies with missing values are different from companies with observed values, and that is reflected in average returns. Now as a second demonstration of a selection bias, I'm going to estimate an asset pricing model. So more specifically, I'm going to estimate the instrumented principal component factor model by Kelly, Pruitt and Su, there's this so-called conditional latent factor model. Now on an intuitive level, what it's going to do is, it's going to take all stocks. Essentially, it regresses all stocks on all characteristics to get characteristics sorted portfolios. And then it essentially applies principal component analysis to this characteristic managed portfolios. And these latent factors are now what I use for my investment. What I'm going to show you here is the out-of-sample of the mean-variance efficient portfolio based on different number of IPCA factors. So it tells me something about the pricing information that the implied pricing kernel of IPCA captures. So I'll show you the result in- and out-of-sample. So what I'm going to do is I take my full data set. I'll use the first part of my sample to estimate the results in-sample. Then I'll use the second half of my sample to show out-of-sample results. So given the estimated loadings of the IPCA model and the mean-variance efficient weights, I'll show you how the results look on the out-of-sample period. So here, the orange bar shows you the out-of-sample Sharpe ratios if I use only the stocks that have fully observed characteristics. The blue lines show you the results where we take all stocks, and I use my local backward cross-sectional model for imputing values. Now out-of-sample, we get substantially higher Sharpe ratios if you include all stocks. It can be almost up to a factor of 2 higher. And the results don't depend on how many IPCA factors, right? The results are uniform with that number. So I get substantially better investment results if I do not throw away data. Now here, on the previous slide, I've compared using either stocks that have everything observed or I -- that means I have all my 45 characteristics observed or I impute all the data. SO it's a somewhat more extreme comparison. Now the question here is, what if I want to include stocks that have 1 or 2 characteristics observed or 3 or 4? I want to have a more refined analysis. And again, you need a lot of these results if you run an investment strategy that combines multiple signals at the same time. Similarly, if you're asset pricing model that should depend jointly on multiple characteristics. What I will study now is I will look at book-to-market ratios. And I will look at simple decile portfolios, namely the top and the bottom decile, value-weighted based on book-to-market ratio. And that will vary which stocks I've put into these deciles. So either we'll put only stocks in there that have book-to-market ratio observed. Or I will only -- I will also require stocks to have an additional size observed or investment on top of that observed or cash flow-to-price ratios or net investments, et cetera, to require more and more characteristics to be observed at the same time, right? So what I'm trying to understand is, what is the effect of requiring multiple signals? Now I use New York Stock Exchange breakpoints for my decile cut-off values. And of course, if I require more and more characteristics to be observed, I would include less and less stocks in my decile portfolios. That's what you can see in the plot here in the right upper corner. Now what is the result of the mean return? So the top decile, that means high book-to-market ratio. Well, if you require more and more characteristics to be observed, actually, the return, the average return of this decile can go up. If you look at the bottom decile, if you require more and more characteristics to be observed, the average return will go down. The reason is data is not missing completely at random, but they are very complex missing patterns. There are very complex ways of how missing -- stocks with missing values are different from those with observed values. That means the effect on mean returns can be either up or down. Now if you have more -- if you exclude more and more stocks, your portfolios are obviously less well-diversified. So what we would expect is that the volatility goes up. That's exactly the case for most of these decile of sorted portfolios. So volatility will go up if you require more and more characteristics to be observed at the same time. What we can see in terms of Sharpe ratio for the top- and bottom-deciled portfolio, with more characteristics, Sharpe goes down because the higher volatility, combined with the effect of the mean return, can lead to this effect. Now this was 1 example with book-to-market ratio. I can show you the same for other characteristics. Here, it's operating profitability and require more and more characteristics to be observed for the stocks there, including my decile sorts. And what you can see is the effect on the mean return can be either way because they are complex dependency effects. Volatility typically goes up, but it can also go down because, again, data is not missing completely at random. But what we observe is once we require sufficiently many characteristics, Sharpe ratios are going down. Now these were 2 examples. Now we show the results for these univariate sorts, that means for the top and bottom deciles for all the different characteristic values. Now what you can see here in blue are the bottom deciles, and in green, the top deciles. And light blue or light green, if I take only the subset of data that has everything observed, while the darker color, dark blue or dark green is if I use all stocks and I use some form of imputation. And for all of these univariate sorts, you see there is an effect. And for most cases, the Sharpe ratios for these univariate sorts will be higher if you use all stocks and you do imputation. And we have the results for all the different characteristics. And similar type of results can also hold for long-short factor. So this was about the effect of the selection bias. Now next, I want to talk about what happens if you use the median for imputation. Keep in mind, the median is not a good way to deal with missing values. So if I use our backward cross-sectional model for imputation over the median imputation, I would get different asset pricing results. But I want to make a statement about what gives me more precise asset pricing results. So what I'm going to do now is the following. I'm going to take my data set, and I'm going to use this logistic regression model for masking data. I'm using this model because it's going to mask data very close to empirical patterns. Then given the mass data, I use different ways to compute it using on my backward cross-sectional model over the median for imputation. And then I use these characteristics to run a cross-section regression. So regress returns on my characteristic variables. This essentially gives me a characteristic managed factor portfolios. I can look at the mean of this characteristic managed factor portfolios, and that gives me the risk premia that I associate with each of these characteristics. I can also look at the time-series of this characteristic managed factor portfolios. And that tells me if I would like to construct a factor that leverages all the information, the characteristics, what would be the time-series behavior. Now because I have data where I did not do masking and I have data where I did the masking, I have now this baseline comparison, right? So I would compare risk premia estimates with the time to information of this characteristic factor managed portfolios to the case where it does not do any masking. Here, I show you the absolute error in the risk premium estimation from these cross-sectional regressions. Orange is median imputation, blue is our backward cross-sectional model. What you can see is that we have uniformly better estimates for the risk premium, and there can be substantially differences. So median implication can strongly distort the asset pricing results that we estimate. Now this was about the mean. Now what about the correlation of this characteristic factor-mimicking portfolios with the benchmark? So we get extremely high correlations with the true values. So it's close to 100%. While when you use the median imputation, you also get a distorted time series. It means you're not only getting the wrong risk premia estimation, also other moments like the variance, the covariance, and et cetera, they will also be distorted. Let me illustrate this. So here, I'm showing you the time-series of this characteristic managed portfolios, right, when you regress returns on these characteristics. The black -- the blue line is what I will call my baseline truth. That is the value when I would take the data without masking. The orange one is when we use our model for imputation. And black is when we use the median for imputation. So what you can see, number one, our implied factor trends here are very close to the truth. The median times here are distorted. So these are cumulative returns of these factor-mimicking portfolios. And that holds for all the different time-series you see, right, for investment, sales-to-price ratios, right? So you get completely distorted moments if you use median imputed values. What is interesting is you also -- here, I'll show you the result for size. Now when we mask data using this empirical pattern, we would never mask size values because the size of the company is always observed. But when you run a cross-sectional regression on all characteristics and you impute some values wrong, it's also going to affect the asset pricing results for other characteristics that are always fully observed. So median imputation can even distort risk premia estimation for size factors, although size was never missing, right? And these are -- this is demonstrating that it can be a substantial imputation bias if you use the wrong method. So let me wrap up here now. So this paper, we provide a systematic study of the missingness -- of missing firm characteristics. It's a pervasive problem. There's complex and endogenous missingness. And simple solutions do not work. We also provide a novel method to deal with missing data. So we leverage the information with time and cross-section information, and we have a latent factor model in the characteristic space. It automatically deals with a very wide range of dependencies and is valid under very general missing patterns. Just to provide an outlook here. The problem of missing data will become more and more important in the presence of big data and machine learning. Because once we use more and more of these big data sets, there's also more missing data that we need to handle. So we believe how we deal with missing data is also of growing importance for specific new data sets like ESG data, international data, et cetera, and the new tumorous implications for asset pricing and corporate finance. We will provide a publicly available data set of our imputed values and also the code so that you can use it in your own work. And I look forward to your feedback. Thank you very much.
Unknown Executive
executiveAll right. Thank you, everybody. We have some time now to answer some of your questions, so we'll go over a few of those. First question we have is what type of missingness is allowed?
Markus Pelger
attendeeThank you very much for this great question. We discussed it in more depth in the paper. But on a very high level, a lot of the stylized facts that we document, namely different characteristics have different probability of missingness. Their patterns -- correlation patterns, either in time or among different characteristics. And the missingness can also depend on being -- on the realization of values. So more extreme values can be more likely to be missing. All of that can be accommodated within our model. Then you guys also point that I think is worthwhile to highlight, the main challenge when it comes to data imputation is to ensure that the imputed values are valid if they are complex missing patterns. So if the missing data is systematically different from observed data, how can you ensure that your values are still valid? And we spelled it out precise in the paper what the assumptions are in a lot of the important cases I've captured with our model. Thanks.
Unknown Executive
executiveGreat. Thank you. Okay. Next question. How much does this matter for machine learning return forecasting?
Markus Pelger
attendeeIt's also an excellent question. Thanks. Now in the paper, we have presented an overview of the results that we have. We showed different investment implication or asset pricing implications. What I want to make clear is that 2 dimensions in which -- the way how you deal with missing data can have an impact on investment or return forecasting. One is which data do you include? That means, do you only take fully observed data? Or do you use some form of imputation to use more data? That is what I call a selection bias. That is where the effect is the most pronounced. That means, if you use your favorite machine learning prediction method, let's say, in your network to predict stock returns, you will get quite substantially different results if you use more data compared to less data. And based on investment metrics, you get better out-of-sample or more out-of-sample profitability if you use more stocks. So I've shown you 1 example with IPCA, but we have done -- we haven't included everything in the paper, but we have done other analyses as well, where we use machine learning prediction of returns. And their results are quite similar. The second question is, how much does it matter which method I use for imputation? Now a somewhat more subtle point, it matters, definitely, if you want to understand, for example, which variable is important. So let's think, for simplicity, about regression. You request returns on characteristics. I use a simple model just for illustrative purposes. Now the way how I do imputation will have a very strong effect on which coefficients will be large or small in this regression. So we want to understand the importance of a variable, the structure of the model. If I use wrongly imputed values, it will distort the model. Now the question is, how is my prediction as well for returns or my investment performance of prediction-based portfolios going to change? This is a more solid question and depends more on the application at hand. Because if imputed values are not extremely -- are not relevant for the return prediction itself, the effect might not be that pronounced. As a general rule of thumb, if you have a better imputed value, it usually leads to a better out-of-sample results. But the degree of -- or in what metrics it leads to better results, that is a more solid question.
Unknown Executive
executiveGreat. Thank you for that. A couple more. One I'll ask is about ESG. So can you talk about how this idea is applied to ESG data?
Markus Pelger
attendeeI also like to ask this question a lot. So for ESG data, we have a quite similar structure as for firm fundamentals. It's 3-dimensional. Now we have a lot of different variables in this ESG data, like admission-related variables, et cetera. SO we're like many variables. We have many stocks, and we observed them a lot time. So the structure of the data is similar. The dependencies are also similar in the sense that a lot of ESG variables are highly persistent. So you want to have a time-series model. There's a lot of dependency among different variables, so you want to capture contemporaneous cross-sectional information. So from that perspective, it's similar. But what it makes it potentially even more interesting is that you have more missing data when it comes to ESG data compared to firm fundamentals. So we think, in that sense, what we document about missing data will be even more pronounced for ESG data, and it becomes even more important of how you deal with that. But overall, our insights will carry over. And the tools that we develop could also be used in that context.
Unknown Executive
executiveThat makes sense. Thanks for that. Okay. Another question here. Did you look at how the forward -- so imputing performed in the asset pricing analysis?
Markus Pelger
attendeeOkay. That's also a very interesting question. So we did not use a model that uses future information for the asset pricing and others where you essentially want to know how well can you trade in the future given prior information, right? We wanted to avoid this look-ahead bias. It would be interesting to see how much the results would be impacted by using future data and if that would give you some infeasible better performance. We haven't done it yet, but it would be interesting to look into that.
Unknown Executive
executiveGreat. Okay. This is a quick one. Will the data set and code be shared?
Markus Pelger
attendeeYes. Absolutely. So we will have a publicly available data set with the raw values and the value without imputation and then the -- our imputed values. And we also will share the code that we used as part of the [ fuse ] for imputation. And so researchers and practitioners can use it for their own data and don't need to deal with the problem of imputation themselves.
Unknown Executive
executiveSounds good. Okay. We'll probably have time for 2 more. Okay. This is the question. Is there a statistical theory for the method?
Markus Pelger
attendeeYes. So this paper or this presentation was mainly focused on the empirical side. We do have another paper that we -- it's a Journal of Econometrics paper that focuses on the statistical theory that is underlying the key element of our message. So we can show formally what are the assumptions on which this is valid, et cetera. And the nice thing is the results would be valid under quite general assumptions. So it's good supportive evidence if you have good theory for what you are doing and in addition to the empirical results.
Unknown Executive
executiveSounds good. Okay. And 1 more. You might have covered it, but does this matter for investment?
Markus Pelger
attendeeI think it does. I mean, I've presented results in that respect. And again, I want to just highlight the different dimensions. I mean, we talked about how things matter. There are different effects, like there's selection bias effect. There's an imputation bias effect. There are different metrics in terms of, out of example, investment performance can be better if you use more data. So throwing away data is not a good idea. How use more data also has an effect. But then you become a little more subtle on what's the nature of the effect is. But on a high level, using a good model for imputation, using more data, will give you better investment results.
Unknown Executive
executiveSounds great. Well, I know we're just about at time now. So Markus, any closing thoughts or ways that people can reach out to you to contact you?
Markus Pelger
attendeeSo thanks, everyone, for attending. Please don't hesitate to reach out to me if you have any questions. The paper is available online. I think you have received the link for the paper. The slide should also be available online for this webinar. And as far as I understand, there will even be a recording in case you want to revisit some parts of this presentation. And again, please reach out to me if you have any further questions or you would like to discuss this topic further. Thanks a lot.
Unknown Executive
executiveAwesome. That sounds great. Thanks again. Thank everybody for attending, and this will be available on-demand. If you have any follow-up questions, feel free to reach out to Markus. Thanks, and enjoy the rest of your day.
Read the full transcript via the API
You're viewing the first half of this call. Get the complete NVIDIA Corporation transcript — plus 248,000+ transcripts from 12,000+ companies, speaker segments, AI summaries and full-text search — through the EarningsCalls.dev API.
Get the API View API docs →This call discussed
For developers and AI pipelines
Programmatic access to NVIDIA Corporation earnings transcripts and 248,000+ others is available through the
EarningsCalls.dev REST API. Plans from $24.99/month — full transcripts, speaker segments,
full-text search, and the recently-added /api/v1/transcripts/recent polling endpoint for ETL pipelines.