In June, Bridgewater ran the toughest public test yet of AI on real investment decisions.
The result: not even close to reliable. Top AI models got it right about half the time, basically a coin flip. Better prompts pushed that into the mid-70s. Then progress stopped. No model beat 80%, the bar Bridgewater's own team said they needed to trust it.
Why the plateau matters more than the score
One score is easy to shrug off. But this team spent months trying to beat it, using every trick available. The ceiling wasn't a lack of effort. It was structural: you can't prompt your way past it.
Where the real gains came from
The fix wasn't a bigger model. It was a private one, trained on the firm's own verified records, that hit 84.7% accuracy, at 14 times lower cost.
Even that model failed at first. The training data had been labeled by outside vendors, and a lot of those labels were wrong. The fix wasn't more computing power. It was a feedback loop: disputed records went back to human experts, got corrected, then retrained the model.
Accuracy is downstream of the archive.
What this means for you
The real finding, buried in the report: how accurate your AI is depends on how verified your records are. A big model reading unverified data will always plateau. A small model trained on verified, sealed records will beat it, for less money.