Language models are poor traders and useful analysts
Given execution authority they underperform. Given a research task they save hours. The distinction is not subtle.
SpansAISECFIN
10 minAlgorithmic & AI Trading
We reviewed a year of results from desks that gave language models varying degrees of authority. The pattern is consistent enough to state simply: the closer the model got to the order, the worse the outcome.
This is not a capability ceiling. It is a mismatch of task shape. Trading requires calibrated confidence under distribution shift and a hard stop on tail risk. Research requires reading a great deal quickly and summarising faithfully. Current systems are markedly better at the second.
Where they earned their place
- Filing and summarising disclosures faster than an analyst can open them.
- Flagging inconsistencies between a filing and prior guidance for a human to check.
- Drafting the first version of a research note that a person then corrects.
- Monitoring news flow against a watchlist and explaining why an item matters.
It reads ten thousand pages a night and it is wrong about position sizing in a way that is expensive. We use it for the first part.
What to watch
Whether calibration improves faster than capability. Confidence that means something is worth more here than another point of accuracy.