Compare AI models

Two models, one window, one asset class, one horizon. Accuracy is shown with a 95% confidence interval, because a rate without one cannot be compared to another rate.

Rule-Based Analysis

rule-based-analysis
33.9%
95% CI 33.4% – 34.4%
  • Sample size30671
  • Correct10392
  • Mean stated confidence 36.4%
  • Confidence gap 2.5 pts
  • Symbols covered 342 / 365
  • Coverage of the field 93.7%
  • Horizons used2
Calibration — stated confidence against measured accuracy
Confidence n Accuracy
0–50% 26079 32.9%
50–70% 3161 42.2%
70–85% 846 28.5%
85–100% 585 38.1%

FinBERT

huggingface-prosusai/finbert
56.1%
95% CI 54.1% – 58.1%
  • Sample size2384
  • Correct1337
  • Mean stated confidence 94.0%
  • Confidence gap 37.9 pts
  • Symbols covered 29 / 365
  • Coverage of the field 7.9%
  • Horizons used1
Calibration — stated confidence against measured accuracy
Confidence n Accuracy
0–50% 0
50–70% 0
70–85% 0
85–100% 2384 56.1%

Over the last 30 days the two confidence intervals do not overlap, so this sample does separate the models.

Accuracy counts a prediction correct when the direction it stated matches the direction the price moved past a 1% threshold over the stated horizon. Coverage is the share of symbols scored in this window that the model expressed an opinion on — a high rate over four symbols is not the same claim as the same rate over ninety. Full method and disclosures