David Holzmüller: hi Ravid. Hi Ellen. thanks for the invitation.
Ravid Shwartz-Ziv: So, we already talked several times on Tableau Data in the podcast, but I think you can give us maybe one of the most, I don't know, state of the art, or what we actually currently are doing with Tableau Data and Tableau Models. So, first of all, tell us, what do you think is the major change? that actually happened in the last, let's say, two, three, four years in the film.
David Holzmüller: Yeah, so I mean I guess the main change, at least in in the predictive tabular modeling field, is these Tab PFN style tabular foundation models. so I mean if you talk about tabular foundation models the There could be like you you could try to pre-train for multiple things, but these models pre-train for learning the learning process itself. So they they take entire tables as as an input. Basically you have the X-train, X test, and Y sorry, and X train, Y train, and X test. And you want to predict Y test, and you just pass it through some transformer model or something like that, and then it gives you the predictions. And so you pre-train them on on data sets, and that's how they how they learn how to learn from new data sets. So I would say that is the biggest thing. And then there were, of course, other development, like basically the improvements in classical like from scratch deep learning were a bit shadowed by this, but there are also advances in this. And yeah, I mean other things about about world knowledge. Yeah.
Ravid Shwartz-Ziv: So, so, so, but, okay, so like now, now I have a company, let's say like medium-sized company, I have a table of data, and I don't know, it can be like whatever, like medical or like your favorite table of data. What I should do? What are the models that I should actually like try first?
David Holzmüller: So I would say the generic answer is you go to tabarina.ai and look at the current leaderboard and this way my answer will be valid in the future as well.
Ravid Shwartz-Ziv: So so start like maybe tell us what does it mean? What is like the Tabarena? Yeah, like that you actually like made how much like two years ago?
David Holzmüller: Yeah, so yeah. So I mean I'm not like the main author of Taburina, so I should give credit to Leonard and
Ravid Shwartz-Ziv: one.
David Holzmüller: Nick and Andre and the people that did a lot of work on that. But yeah, Taburina is a benchmark that tries to be the best benchmark for tabular data. So there are some aspects like spending a lot of time on going through dataset sources and making sure to pick realistic data sets with like from realistic predictive tasks that don't have leakage and these kind of things. And then running running good baselines, which means like, for example, my part in Taburina, or like maybe my main part in Tab Arena was making sure that the boosted trees and random forest baselines were had good hyperparameter tuning spaces. so I like used some handwritten tooling and stuff to to optimize these search spaces just to make sure they're they're they have very good search spaces and then like we set the high tree budget and we tune them using a lot of random search and then we use this also Karana style greedy ensembling like that's with cross-validation and and everything to make sure that these these are really well tuned And like the the automatic preprocessing from AutoClone. and then and then yeah, it's also a living benchmark, so it has a live leaderboard. It's being updated with new models. so you can go there and look at different data set subsets and metrics to see which kind of models are the best ones for your kind of use case. I I mean like to to get an idea of
Ravid Shwartz-Ziv: So...
David Holzmüller: that.
Ravid Shwartz-Ziv: Okay, so... So first of all, like, okay, so let's talk a bit more about Taberian and then, like, we'll go back, like, who... what is
David Holzmüller: Mm.
Ravid Shwartz-Ziv: the best model, maybe, but, like, what... what was the problem? Let's start from the beginning. What was the problem with previous, like, benchmarks? And did you decide to actually, like... or do you think that, like, Taberian are solved?
David Holzmüller: well there are a few I mean the data set quality maybe. Like I think It's a bit unclear how much of a problem this was, but we know there were some like there's also some research from like Yandex, TabRed, and whatever, they found that bunch of data sets were like temporal, and if you just do random splits on them, you actually favor methods that can look up very nearby points because they're available in the random splits. So you have like these retrieval-based methods that are very good on these data sets completely. Compared to all other methods. Like I saw methods that had like an RMSE of 0.06 instead of 0.36 or so. And that's maybe not what you want. And then and then yeah, the the tuning of the baselines, but I think this has also gotten better and better already over the years, but still. And then a lot of benchmarks. Also for just compute reasons, they don't have the they don't run methods with inner cross-validation. Whereas we like we run them with eight-fold cross-validation and then ensemble all the eight trained models for the classical models. So that then you can see, like we saw, for example, in the paper, if you ensemble MLPs. you get more boost than if you ensemble boosted trees because they are already ensembles basically. and and yeah, just yeah. And and and then also like this being a living benchmark that is updated and has a leaderboard so in so that next year it's hopefully still up to date.
Ravid Shwartz-Ziv: But you're not afraid that we'll see the same thing that we actually seen with LLMs benchmarks, that basically they are saturated quite fast and it's very debatable if it's something that's related to the overfitting of the models or like these new capabilities.
David Holzmüller: So with saturation, I mean People have said that tabular data is quite saturated. Now we see actually a lot of progress. I don't know when it's going to be saturated, but not yet, I think. I guess I guess there is one concern of people like training on the benchmark or overfitting to the benchmark, which I mean if they don't disclose disclose it, we can We can't do anything except add new data sets gradually over time that will hopefully expose this. So now we have a Beyond Arena, which is basically the next generation of Taburina that will be integrated back into Taburina. that has more data sets, more types of data sets where we have like datasets with temporal splits, with grouped splits, larger and smaller and these kinds of things. And then then we can see if some methods are like fundamentally worse than before.
Allen Roush: So what what I'm curious about because I the last time I I really looked at tabular data was, you know, twenty twenty two around this. Two two questions. One, you know, I remember profit back in the day getting in trouble for making claims about its performance that didn't pan out. And two, somewhat related to this, there are GPU powered Auto ML solutions today that are mostly r themselves running, you know, maybe QML, like the NVIDIA rapids implementations of many of the algorithms. And do you think that it makes sense on many problems to just run GPU-powered Auto ML? Or do you think that these Tab FN like foundation models are so much better than than trying to recruit from first principles?
David Holzmüller: I so I think the best thing you can do is or I guess it depends. Nowadays I don't know that much about other AutoML tools. I know that Autogluon integrates these tabular foundation models to stay on top, but the the foundation models come out so fast that sometimes they're not yet integrated and then sometimes they beat Autogluon and then Autogluon integrates them. So that's like a a cycle, but yeah. so if you want to like use a GPU and spend a bit of compute then like AutoGluon can tune over these different models. Otherwise if you have a CPU then AutoGluon can tune over CPU based models. yeah. So yeah. I I guess it's like still if you wanna get the max performance then you should still like do a tuning and assembling of different models.
Ravid Shwartz-Ziv: Okay, so let's assume that now we have good enough benchmarks. What is like now, yeah, so how I can, what do you think, like how I choose the best model? Just go to a tab arena and choose the one that is the best? Let's open and let's see if now I'm looking. So,
David Holzmüller: Mm-hmm.
Ravid Shwartz-Ziv: tabfm is like the first, the best one. Actually, I
David Holzmüller: Yeah.
Ravid Shwartz-Ziv: thought it like tabpfm, but actually, now.
David Holzmüller: Yeah, that was like what, four weeks ago or three weeks ago, Tab FM came out. And yeah. So of course you should also look at the trade-offs. Like Tab FM is a huge model, it has like one point five billion parameters. and it is I don't know, ten, twenty times slower than the other like Tab Fn or Tabical. So it always depends on like your constraints and your data set size. If your data set is much larger than one hundred thousand samples, probably you want to use a classical method because it will be better than that. Then the
Ravid Shwartz-Ziv: A classical method is like writing boosting trees.
David Holzmüller: Gradient booster trees and MLPs. I mean for the MLPs it depends how much software support you need, I guess, because they're not as mature as the gradient booster trees. And also they the MLPs will still need a GPU, but they do scale to quite large datasets. like we're also seeing on Kaggle Like they have these tabular playground competitions now that for the last maybe year or so have been basically always data sets that have like six hundred thousand samples. And there is always like like MLPs or booster trees at the top.
Ravid Shwartz-Ziv: But not like, I don't know, Tab FM or something like that.
David Holzmüller: No, not yet. I mean Tab FM is probably also
Ravid Shwartz-Ziv: or like tab and right, like there are like a lot of.
David Holzmüller: hard to run on Kaggle GPUs, I would imagine, but maybe it gets close nowadays.
Allen Roush: W what about what about things like the loss functions when they're they're training and what about like is cross-validation still being used? Like
David Holzmüller: So you mean for for the foundation models?
Allen Roush: No, no no, I mean even these Kaggle comp t competitions. And when I mean loss functions, I mean like, you know, log loss versus KL divergence versus standard accuracy or F1 or I remember I had to make all these choices back in the day when I cared about tabular data a lot. And I I don't know if like especially for things like calibrated confidence in the answers, right? Which some people really care about. I'm just like wondering what
David Holzmüller: Yeah.
Allen Roush: developments there have been.
David Holzmüller: So I mean for for for the training loss, I guess everyone uses log loss normally. and then for like early stopping cross validation you would normally use whatever is the downstream loss. There isn't Like in Kaggle competitions they rarely do like log loss or Briar score as like a downstream metric. So unfortunately we don't see many people trying calibration things on the models just because it's not necessary. I mean in the best case it's like F1 or balance accuracy, you have to do some threshold tuning. yeah. We we've been looking in basically my Other line of work on like uncertainty quantification a bit and into calibration and things like that. And so we we've seen like you should you should integrate the the po like postal calibration methods and you should early stop on maybe on the lost after postal calibration instead of the lost before post-calibration. But yeah. Still have to wait until there is a Kaggle competition to demonstrate it, I guess.
Allen Roush: And and is there value still in in I guess clustering or over under sampling or class and sample weighting, like these kind of methods?
David Holzmüller: I'm I'm not an expert on that. I mean we have imbalanced data sets in our benchmarks. We don't do any of these things. I hear different opinions from different people, but I haven't tried over or under sampling.
Ravid Shwartz-Ziv: So what do you think? Do you think we'll still have this separation between models that we can run on small, not small, not huge number of samples versus models that we can run on millions of examples? Or do you think we'll see kind of converge?
David Holzmüller: I mean the the tabular foundation models will surely keep scaling more towards large data. I don't fully know how, if there's gonna be some kind of linear attention or context selection or fine-tuning or whatever. But I I think I I don't see I mean with context selection theoretically you could get to like billion scale datasets, but maybe that's not attractive even though you could do it. so I I would expect that there will that there will remain this distinction for at least a while.
Ravid Shwartz-Ziv: So what do you think are like the... Where are the cases actually like foundation models like are better? Like do you think we will... like this is just only a matter of like how much data do you have or do you think that actually there are cases that you can actually generalize from you know other related problems?
David Holzmüller: sorry, could you repeat the question?
Ravid Shwartz-Ziv: Yeah, so we have these foundation models, right? For example, like Google's ones.
David Holzmüller: Mm.
Ravid Shwartz-Ziv: Right, like TABFM. Do you think there are problems that these models are better than specific models? Than, you know, TABFM? Besides the number of samples, do you think there are problems... with some specific properties that we can actually say, okay, these problems are good for tabfm, right?
David Holzmüller: Yeah.
Ravid Shwartz-Ziv: Versus problems that the in-context learning models are better.
David Holzmüller: Yeah, I mean this is a quite hard question in the sense of people have tried to do meta feature analyses to predict from meta features of data sets which model is good. But it seems that we would need quite a lot of data sets to reliably learn this. And it seems that a lot of meta features with the current number of data sets maybe are we can't really say if they have predictive value. So dataset size seems to be one. From Beyond Arena, like w for grouped data, if you have groups and you make group splits, that seems to be another one. Maybe like number of features will be one, but yeah. Otherwise it seems quite difficult to To predict this.
Allen Roush: Ever since Chat GPT came out, have you seen like a reduction in like focus or demand or effort put into like Kaggle submissions or similar like tabular like do you think that there was a serious trade off and reduction in investment in that field?
David Holzmüller: I mean there are There are quite few like prized tabular competitions nowadays. I think maybe if they're like a bit specialized to like tabular plus some special objective, but I'm haven't been on Kaggle long enough to have been in the tabular era, but I've heard it existed in some
Allen Roush: Yeah.
David Holzmüller: some sense but I don't know if like Chat GPT had an that much of an influence on this because I think I mean before people were also looking into into image and and like different modalities. I mean I guess compared to compar compared to the overall field probably tabular is more popular nowadays than five years ago.
Allen Roush: And how important is model explainability to the modern field? Like, and and what techniques are people using? I mean, obviously things like like you know, d variations of decision trees are sort of in many ways explainable by design, but you know, if if MLPs are getting used, then maybe they're using some post hoc technique. Is it all still Lyme or have they gone to far subsequent methods?
David Holzmüller: That's a good question. I mean with all the MLPs and foundation models, I mean most of the time they're treated as black box. I ha I hear much more of Shap than of Lyme. Like TabPFN and Tabical have Shap adaptations. so you can compute Shap values. You could of course throw permutation importances. Yeah, other than that, there is some work in like making making more interpretable models. I've seen work that tries to either directly predict sharp values with the foundation models or tries to make a last layer that like retrieves labels from data points and so you can see basically the attention map of the last layer, like which other training samples is it attending to. Or models that or foundation models that output generalized additive models. And then you can interpr interpret those. You just don't know how they were learned, basically. So yeah, there seems to be something but nothing that or Yeah, if you want max performance you have to go with a like black box model that you can only postdoc interpret.
Ravid Shwartz-Ziv: And what do you think about, like, LLMs in general? So first, why you think, like, LLMs are not good in... tabular data? And then if you think there is a way that we can combine the two, you know, both, like, tabular data models and LLMs.
David Holzmüller: So there are a few issues. One is just like efficiency and context size. If you in an L to put it like if you want to use LMs for in-context learning with tables, you have to serialize the table into into a string. So you have to use tokens from a regular LM vocabulary, so you probably use more than one token per cell. then of course the LMs are mostly pretty large compared to the tabular models. And then they also have this one dimensional attention, which means that you basically grow in complexity like number of samples squared times number of columns squared. Whereas with the tabular models you have this kind of 2D attention and then often like compression At some stage, which gives you more like number of samples squared times number of columns, and then number of columns squared times number of samples, or even better. So like just scaling to larger datasets is a problem. I've seen the the largest scaling I've seen with LLMs is 2048 shot, which is not much more than we had with Tab PFN1. Whereas now in the tabular models we can easily go up to a hundred thousand rows. And then of course you need to pre-train also the LMs for the for the task. So you can do that. Usually like some people like to post train LMs for it. And then the question is what is the benefit that you expect? Mostly I guess it's expecting the benefit from world knowledge, like if there are strings in your columns, or if you have column names, or if you maybe you have a description of the dataset, but I then we would also need to have a dataset that's like a pre-training corpus that has these things.
Ravid Shwartz-Ziv: But you don't think that the last point is quite very significant, right? I would expect that if I have information about this data, Like, variable information about this data, it will be much easier to solve it, right?
David Holzmüller: I guess it depends a lot on the size of the data sets. I I would expect that once you have like a thousand or ten thousand rows, it might be very hard to come up with a description that is not already learnable from the data itself. Like if If you describe that the f target is periodic in this feature, it's probably recognizable from the data itself.
Ravid Shwartz-Ziv: But you don't think that like...
David Holzmüller: But if you have small data then I would expect it to be useful, of course.
Ravid Shwartz-Ziv: if you have small data.
David Holzmüller: Yeah.
Ravid Shwartz-Ziv: Yeah. So, like, why we don't say it? Like, why we don't say that, like, even on small data, like, LLMs are not good as, like, other models?
David Holzmüller: I well, I the thing is I think we need better benchmarks on small data to figure this out. Because of because
Ravid Shwartz-Ziv: So you've seen.
David Holzmüller: of the context size limitations. I mean we we need benchmarks that are on the the sizes that the LMs are actually good at. So I mean with Beyond Arena we go down to one hundred samples, but we don't have LLMs on Beyond Arena. But if we could like make a benchmark that goes down to these sizes that actually figures out what are good baselines for these sizes of data sets, at least for like boosted trees we would probably need to but retune the hyperparameter space and these kind of things.
Ravid Shwartz-Ziv: But do you believe that for 100 examples, if I will just insert it to GPT or Fable, do you think they will be better than both classical models, like CatBus or even the recent ones?
David Holzmüller: So we we tried this for tabicol v2 for rebuttal. with eight samples opus four point six was better than tabicol, but with sixteen samples tabicol was better than opus. So I mean probably if you would post-train these models on like really tabular tasks, then you could get to something and also if you had data sets that have actual text in them and textual descriptions. 'cause I don't think we had that, but yeah, then we would need a benchmark that has such data sets. So I
Ravid Shwartz-Ziv: And.
David Holzmüller: I really don't know r realistically when LMs are good on tabular data or not. And then you also have the whole
Ravid Shwartz-Ziv: But do you like it?
David Holzmüller: contem data contamination issue that you need to solve.
Ravid Shwartz-Ziv: But it's quite surprising, right, that like, no one actually like... try to rigorously evaluate it, right? Because people are using LLMs for everything these days. And tabular data is such a huge field, right? You have so many problems and so many data sets and so many places are dealing with tabular data. So it'd be so much easier if instead of calling for the latest coding problem. will just call it with my latest table of data problem.
David Holzmüller: I mean, would it be much easier than having a tool call a Tab Pfn or Tabical or whatever? and and then you also get something that is much faster as well. So I don't know if these models should be the ones that also solve the tabular problems.
Allen Roush: do you do you think there's room for creating like whole new classifiers, for example? like like i have we panned out all of the methods for classification or regression that could exist in your mind, or is there like a whole new world potentially out there?
David Holzmüller: That's a good question. I mean there could be. I mean we we kind of ventured a bit into this with XRFM, which builds on RFM, which is recursive feature machines, which is basically kernel methods with learned feature importances, which is like something that is very unusual on tabular benchmarks in the last few years. And we show that it can be like close to boosted trees or maybe better in some settings. so I guess there is some room for it. yeah. And and especially if you if you want to create like efficient methods or or methods for certain with certain constraints, but with I I I I guess it will be quite hard to to beat the scenario where you can just directly pre-train for optimal learning performance. Basically.
Ravid Shwartz-Ziv: Where do think the field is going? Where do you see, like, how do you see, like, models in five years from now? Both models maybe and also, like, the benchmarks.
David Holzmüller: Hmm. So models I I mean I expect them to get better of course. With with Tab Fm especially, the question is also if they will get a lot larger, which can like run counter to the goal of also scale making them scalable to larger data sets. So maybe there will be like different trade-offs and you have a large model for small data or a small model for large data.
Ravid Shwartz-Ziv: Do you think there is no advantage for large models for large data?
David Holzmüller: no, I I I think there would be an advantage. The question is if you want to afford that and if you can fit it into RAM because RAM is also an issue for large data. It's not only it's not only the quadratic runtime of attention. And o of course you could just try to parallelize it across more GPUs, but maybe you don't want to deploy it at that point.
Ravid Shwartz-Ziv: But do you think we will see more foundation models? Or do you think we will see more top-tier fan styles that you train on some random data and then to apply it to specific
David Holzmüller: So
Ravid Shwartz-Ziv: data?
David Holzmüller: I don't know if your question was referring to like real versus synthetic data.
Ravid Shwartz-Ziv: Both. Like, both, like, what is, like, do you think, like, now we will see, like, models that train on, like, different datasets, different types of data, and then you take these models and just apply them to your data? Or do you think we will see, like, models that are very, like, narrow and just train them only on your data?
David Holzmüller: So I mean one thing you could do is if you have several datasets, you could fine-tune them on some of the datasets to make them better on like the one downstream dataset. For the real versus synthetic question, I I don't know what is going to happen. For now the synthetic data seems to Be prevalent and it has like a bunch of advantages, like it's much less risk of contamination. Like you will, of course, you can optimize the synthetic data on real benchmarks and by this way get contamination, but it's not like you can just memorize the the data sets. you can generate arbitrarily many and large data sets and I guess one other thing that they will be very useful for is making causal foundation models. I mean they're already being made and have been made because you cannot really do causal interventions on real data where you don't have the model for the data. And then real data should be much better in all cases where like text or maybe even images are involved, like anything that needs like world knowledge. And this goes a bit back into the involvement of LMs and in tabular data. There's one other way in which you can use LMs, which is for just embedding text in text columns. And this is like we're looking at this now, it has already been looked at and There are some questions on like what is the best way to to embed text for use with tabular foundation models and maybe when you embed the text it gives you something very high dimensional which makes it very slow, so can you accelerate that? yeah, then how do you pre-train
Ravid Shwartz-Ziv: So what is the best way?
David Holzmüller: it maybe directly?
Ravid Shwartz-Ziv: So what is the best way for embedding?
David Holzmüller: So I mean I can talk about the things that are released. so we saw in the strable benchmark that it can make quite a difference in the sense that the first thing we tried is just taking like embedding so for each text cell we run it through an LM and then the LM gives maybe like four thousand dimensional embeddings and we just run a PCA and project them to one 128 or something like that, and and then run the model on that. Or I I think actually we projected it to 30 dimensions and we saw that sometimes this failed spectacularly, especially with the tabular foundation models. and and then if you standardize each like each dimension of the embedding before you run the PCA, it gets quite a bit better. And then there are these LMs that have Matryoshka embeddings. So they're trained to so you can take like the first 32 dimensions or the first 64. and and they're valid embeddings in and of themselves so you can just cut them out and we saw that even for even for LMs that are not trained for this it can be better to to just use a fixed number of dimensions instead of running the PCA. And then there was another paper that also ran standardization across the embedding dimension instead of across the number of samples dimension. And I think there is a bunch more room for how to embed values. I saw a paper that said like you take a post-trained LLM, like post-trained for tabular data, and then you use 128 rows of the data as context, you you cache them, and then you use your new row, and then you embed your new row with the like You take the embeddings of the new row and then you pool them. so I think there's still yeah there are a bunch of ways to do this more cleverly.
Ravid Shwartz-Ziv: And do you think at the end we will see a unified embedding?
David Holzmüller: A unified embedding
Ravid Shwartz-Ziv: for like both across like different modalities and also like for even for tabular data.
David Holzmüller: Well I guess one question is whether I mean one question is whether you want to have contextualized embeddings or not. Like if you have a data set, do you want each row to have an embedding that is independent of the other rows or not? If you want it independent, you can't use a tabular foundation model. and then the question is if there can be a an embedding that is sort of good invariantly of the downstream task, which I don't know.
Ravid Shwartz-Ziv: So maybe, the answer is maybe.
David Holzmüller: Maybe, yeah. I'm I'm not I'm not too much in the embedding research to have a strong opinion of that.
Allen Roush: are you familiar with like time series and regression stuff?
David Holzmüller: Yeah, I'm I'm I've been looking a bit into time series forecasting stuff, like trying to follow a bit what's happening.
Allen Roush: H are they also doing similar stuff with foundation models having taken over?
David Holzmüller: definitely I've I've heard from some people that apparently the foundation models are not as good as the like as in tabular. Maybe also the baselines and the benchmarks are not as good. Because at least on the benchmarks they look very good. But
Ravid Shwartz-Ziv: You
David Holzmüller: yeah.
Ravid Shwartz-Ziv: Do you think there is a fundamental difference between tabular data and time series data?
David Holzmüller: I guess time series is can be much more complicated. because
Ravid Shwartz-Ziv: Complicated. Why?
David Holzmüller: you you can have covariates, which are basically kind of all the complexities of tabular data, like you can ev irrelevant you can have irre irrelevant covariates and different causal relationships between them and whatever. And then you also have all the different
Ravid Shwartz-Ziv: but you can have it also in tabular data, no?
David Holzmüller: kinds of temporal relations on top. But there are also some
Ravid Shwartz-Ziv: But you have it.
David Holzmüller: yeah.
Ravid Shwartz-Ziv: You don't think you can have it also in Tableau data?
David Holzmüller: Temporal relationships.
Ravid Shwartz-Ziv: for example, or like other relationships.
David Holzmüller: Well, I guess at some point it becomes a question of definition if you if you have so like
Ravid Shwartz-Ziv: No, I mean like...
David Holzmüller: at at which point does it become a time series dataset then?
Ravid Shwartz-Ziv: Yeah, I agree. do you see? But in tabular data, you can have a lot of different types of relationship. Some of them are related to time. Some of them are related to other stuff. And you basically need to reveal this information from the data. So as I see it, it's a more complicated case,
David Holzmüller: Yeah, yeah. I mean, there are also so what we have in the Beyond Arena benchmark are temporal tabular data sets. And
Ravid Shwartz-Ziv: Mm-hmm.
David Holzmüller: the distinction we make in the paper is that temporal tabular datasets are ones where you Don't like in in in time series you forecast now and then you observe it in the future what what happens and in temporal tabular like you train now but you make the inference in the future for the present in the future, which is maybe a bit complicated. I I guess another thing we see, for example, when you apply tabular models to time series, like there's Tap NTS, there is the tab table adaptation to time series that is based on Tapheev NTS. If so y you have your history and then you have your future I mean if you have covariates in the time series, you have your future covariates. Like let's say you want to predict the wind power plant generation and you have the weather forecast for a certain amount of time, and maybe The tabular model would make a prediction on like for a certain time point while seeing only the covariates at that time point. Whereas a time series foundation model, depending on whether it's a decoder type of or encoder type model, it would see the history of covariates up till that future point, or it would see even like the covariates after that future point. So it depends a bit what information you think will be available at that point. But certainly like there are attempts to make I mean, tabular models are used on time series and you can make synthetic data that also includes time series things. So
Ravid Shwartz-Ziv: Do you think we will see... Maybe, do you think there are cases that gradient boosting trees are better? Like, it's better to use them? Like, they are easier? Like, there is any advantage of them? Or like, you just say, don't look on them, look on the... don't know. Whatever is like the first on top arena and go from there.
David Holzmüller: no, there are definitely use cases for gradient boosted trees. One is larger data sets, one is you want to have a model that is fast on a CPU. you want to have a model that has a fast inference time. Maybe you want to use them as a even if you use a tabular foundation model, you want to use them as a distillation target to get the fast inference time. You can compute chap values much faster for booster trees, as far as I know. you have a bunch of things that you can put in there like monotonicity constraints or Like you can maybe control them much more directly. Maybe you'd yeah. I mean people still use linear regression as well, right? So I don't think they're that these these
Ravid Shwartz-Ziv: You
David Holzmüller: methods are going to disappear. And yeah. If of course if you look at Taburina, the benchmark goes up to one hundred thousand samples right now. So if you have a data set that is much like outside these constraints, yeah, maybe you Maybe you s use other methods.
Ravid Shwartz-Ziv: And what about new models? you think, like... So there was a point that we had a lot of models, like new models, based on gradient boosting trees. And it looks like these days, everything, all the new models are based on transformers and things that are related to these. Do you think there is still a place for improvement for... running boosting trees or... that's it, this is like the best that we can get.
David Holzmüller: I think there is a I think there is some room for improvements. So one thing is that I mean there are like the three big ones, right? CatBoost, Lite GBM and XG Boost, and then there's also the CikitLearn one. And not all of them have the same hyperparameters. So I guess if some of them would adopt the hyperparameters of the other one, they could get better. I th I don't expect this to be like a huge difference, but at least a bit. But it seems that like the the development of gradient booster trees is not so active. So maybe this will not happen in a while. then I mean one big thing is feature engineering that I haven't talked about, but of course if you want to know and and that is like one one big thing that we don't have in the benchmarks because it's kind of hard to to do, is to to know which method actually works best after you do feature engineering. and maybe that's where the boosted trees catch up. so there is a paper by some of my co-authors called Tab Prep, where they basically make an automated feature engineering pipeline and they show that especially for like light GBM this can improve the results a lot on on larger datasets. So but interestingly, I mean It seems from their results that the improvements of this feature engineering are more or less orthogonal to the improvements through TFMs. In the sense that they are not on the same data set. So it's not like TFMs, so like tabular foundation models only give you the same improvements that you would have gotten through feature engineering otherwise. So yeah.
Allen Roush: So
Ravid Shwartz-Ziv: And so, for future engineers, how much do think we actually need? How to say, maybe, expert knowledge or prior knowledge about these data sets? Do you feel that this is important these days or not anymore?
David Holzmüller: I mean, I guess again I would say that probably depends a lot on the size of the data sets. If you have a small data set you need to guide the method towards the right things because it cannot reliably figure them out from cross validation. If you've larger data sets, well I mean the problem is that we don't really have a good way to measure it in an automated way. Basically we would need to hire a Kaggle Grandmaster, let them do stuff on a lot of data sets and compare that to some automated solutions and yeah. Or or I guess we would we would need to run actual Kaggle competitions, having data sets where the test set is not available online.
Ravid Shwartz-Ziv: I love it, like you've seen it, like we should do it in a rigorous way. In
David Holzmüller: Yeah.
Ravid Shwartz-Ziv: LLMs, people like, there are so many companies that are talking, for example, about self-improvement, and they are not even thinking about making some rigorous comparison, right?
David Holzmüller: Mm-hmm.
Ravid Shwartz-Ziv: So, no, that's good. That's good.
Allen Roush: Yeah
David Holzmüller: Yeah.
Allen Roush: y' so I'm curious, so we have all these you know, now tab tabular foundation models, and gradient boosted trees are still being used. Has anybody ever tried to make a language model with any of the classical methods, like GPU powered versions of them? Like what's stopping us from a very large gradient boosted forest t being used on like a tokenized vocabulary to be an LL? And I I'm sure obviously we need to figure out ways to like implement temperature and and all of these, right? Like without the standard softmax, it's but like has has have you ever even seen an attempt like this or even think it's possible?
David Holzmüller: I'm n I'm not sure what what you envision. Like how how do you make a gradient booster tree into an LM?
Allen Roush: Like classical methods to generate it. Yeah, yeah, exactly. Like like why can't we c convert these into generative methods?
David Holzmüller: Mm. So you g I mean generative like you could you could train them to predicts like the next cell in the table or something like that. well I guess one I guess one issue is that like at least naively if you want to learn them to predict a lot of targets, like let's say you want to learn them to predict all columns in the table and not just the target one. just the runtime increases proportional to the number of targets, 'cause there is no like shared network, like in in neural nets, that has like the same cost. Although they're
Allen Roush: Well well this happens with neural nets with vocabularies too, right? Like the vocabulary going up on an LLM actually increases the computational cost.
David Holzmüller: Sure, but then you have all these middle layers that don't really d depend on the vac vocabulary size. So it doesn't scale as badly with the vocabulary size as like if you make one tree for each targets. But yeah.
Allen Roush: Yeah, makes sense.
Ravid Shwartz-Ziv: So what do you think are the missing components? Or the thing that's worth to do research about them, or to try to develop in tabular models?
David Holzmüller: Mm. I think there is a lot of things. So maybe what I didn't mention previously is of course there will be like scaling and this will be maybe like the one of the biggest things for like industry labs, like scaling to larger datasets. But yeah, then you could look at things like the current architectures are not invariant to class encoding or column order like and and it's it's like not easy to make them I mean there are some approaches but they haven't been that popular so the question is can you make a good model that is invariant to these things? how how do you develop synthetic data generators? you can try to make tabular foundation models that don't just try to do classification and regression. And there's of course some work on this, like survival analysis, causal predictions, some uncertainties, there's a connection to To Bayesian machine learning, like how close are they to being Bayesian? And can we can we measure it? Can we get uncertainty decompositions? what happens to their calibration? How can we distill them into other models? How can we combine real with synthetic data? Can we make these models even faster, like scaling them down instead of up? Or I I think there can be a lot of stuff that is yeah.
Ravid Shwartz-Ziv: So do you think you will see more research? know, like, of course, like, there are a lot of people that are working on tabular data, but it's still, think, like, quite small compared to language models, right? Coding agents and things like that. Do you think you will see more research on tabular data or...?
David Holzmüller: Yeah, I think it's
Ravid Shwartz-Ziv: the bigger problems to solve there?
David Holzmüller: Yeah, I yeah, I think it's it's growing a lot. Like the workshop participant numbers are basically doubling each year for the tabular and time series workshops. I guess partially like people are jumping on board from like different communities, like uncertainty or Bayesian models or whatever, and then probably also just in general more people finding this interesting than before. And now I mean we have I mean we've seen like Google enter the race a month ago with Tab FM. So I think that will make it even more popular. And I mean you you also see the growth of the startups like all of them being bought or valued at a billion dollars or and probably growing even more so I think yeah. It's going to grow a lot more.
Allen Roush: R related related to that, the stocks for tech stocks have been doing really bad recently. Do you think AI's in a bubble and do you think it's popping?
David Holzmüller: I I I didn't even look at the text talk, so I didn't know that. But I mean I can Yeah, I guess AI is AI might be in a I AI might be overvalued or not, I don't know. there is
Ravid Shwartz-Ziv: You
David Holzmüller: definitely some value to it. And yeah I I I find it hard to say this.
Ravid Shwartz-Ziv: I will ask it differently. Do you think there is something from LLMs, you know, it can be like more technical stuff or like maybe more general like the processes and how we are like dealing with problems there that we can use for tabular data, tabular models?
David Holzmüller: Yeah, I mean definitely since they're also transformer models, all the transformer advancements are potentially interesting. There is a bit of like a challenge in the sense that if you if you want to have like linear attention or things like these, like the LLM ones are usually designed for ordered sequences and tabular data is not ordered. So in this is like maybe the one thing where time series actually has it easier than tabular because time series are ordered. Yeah. I'm probably I mean, I guess there is a question of can you make reasoning models for tabular data and what does that even mean? Which I can't answer you. But yeah, I mean probably there is a lot of stuff to be taken from the LM community. Yeah.
Ravid Shwartz-Ziv: Okay, we are almost out of time. Do you have anything else that you want to add? that you want to talk
David Holzmüller: I mean
Ravid Shwartz-Ziv: about, you want to advertise, you want to sell.
David Holzmüller: Yeah, I mean I I I I would find it interesting also to look into better MLPs, like real MLP. Now there is also developments with Tab Pack at ICML to make faster MLPs and these kind of things. And I I hope there will be also advancements with MLPs, at least for the large datasets. yeah, unfortunately I can't really work on everything at the same time, so I I I don't work on real MLP at the moment, but well.
Ravid Shwartz-Ziv: It reminds me, so, like, in real MLP and, like, this line of work, like, you put a lot of effort in order, like, to optimize the things, right? Like, the hyperparameters and to tune them. How much, like, is... it's a problem, like, in current time, you know? Like, how much is a problem like that? The bass models are not, like, well optimized and well tuned.
David Holzmüller: Mm, I guess it depends. Like if you go to Boosted Tree libraries, like if you go to XG Boosts and you just use it with default parameters out of the box, it will be quite bad because the default parameters are set for just having something quick and maybe also not Like there's no cross-validation, no early stopping, and maybe they were set quite early and now they're have to be backward compatible. Whereas for TFMs you just load the next checkpoint and the library's going to be updated to use the newest checkpoint. Yeah, so I but I I think That's maybe something that like classical machine learning has suffered from. Not not giving you like a good thing that works out of the box. And yeah, I mean Catboost is quite good at that, but other than that yeah.
Ravid Shwartz-Ziv: Okay. David, thank you so much for joining us. It was a great pleasure to have you. And I hope you will solve all the tabular data sets that exist in the world. In one model.
David Holzmüller: Yeah, it was a pleasure to be
Allen Roush: Yeah, it's a pleasure to chat with you, David.
David Holzmüller: Yeah.
Ravid Shwartz-Ziv: Thank you.
David Holzmüller: Thank you.
Ravid Shwartz-Ziv: Okay.