Hacker News new | past | comments | ask | show | jobs | submit

Why Large Language Models Fail at Tabular Prediction

https://arxiv.org/abs/2608.02412
The first thing I'd do if working with an LLM on tabular data is to ask what the best tool would be to work with that data and build up a proper harness to work with the data sensibly. Rawdogging LLM isn't the tool for forecasting like this, as they found.
loading story #49170360
Look at the white text on white background in Appendix F. Pretty funny.
loading story #49169933
loading story #49170386
loading story #49171382
Unsure if it's LLMs that fail at tabular data or its just that tree boosting are spectacular at that task.
{"deleted":true,"id":49168861,"parent":49168128,"time":1785850818,"type":"comment"}
Non-LLM transformers beat tree boosting - TabPFN, TabFM

https://research.google/blog/introducing-tabfm-a-zero-shot-f...

loading story #49170591
One step further are those who want to point an llm directly at the data warehouse to get the data needed to run predictions
loading story #49170879
loading story #49171241
Just have 2 LLMs debate whether tabs or spaces are the superior choice
It has to be 3 in case of a tie. Like the magi system in evangelion.
I don’t know, we saw where that led and I’m not interested in becoming a pool of orange tang yet.
No human souls imprinted in the LLMs (Yet), that I am aware of. I think we're safe for now.

But I'm absolutely joining the Machine Crusade if we have a first Impact event and I survive. Some days I wonder just how flabbergasted Frank Herbert and other pioneers of sci-fi would be at the situation we find ourselves in today.

Nowhere in the paper do they mention the reasoning level or budget used for the experiments?

You’ve got to be kidding me. That one variable could make a huge difference in the results. I can’t understand why they would leave that out.

>We study a frontier LLM in its purest inference regime - a single generation pass over a prompt containing the full training and test data, with no tools, no agentic scaffolding, and no fine-tuning

Sigh. So this is somewhat interesting niche academic research but utterly irrelevant to real-world use cases.

loading story #49171053
I find that an odd take. The paper claims to establish what causes the problem: dimensionality. They are clear in that they don't understand why. But this sort of work is what needs to be done to eventually solve the problem.
Solve what problem? My hammer can't drive screws. Is that a problem to solve?
loading story #49172080
loading story #49170109
loading story #49169429
loading story #49169457