Compare AI models

Two models, one window, one asset class, one horizon. Accuracy is shown with a 95% confidence interval, because a rate without one cannot be compared to another rate.

Rule-Based Analysis

rule-based-analysis
28.9%
95% CI 28.4% – 29.4%
  • Sample size32126
  • Correct9281
  • Mean stated confidence 33.7%
  • Confidence gap 4.8 pts
  • Symbols covered 1050 / 1064
  • Coverage of the field 98.7%
  • Horizons used1
Calibration — stated confidence against measured accuracy
Confidence n Accuracy
0–50% 29851 28.1%
50–70% 1394 42.2%
70–85% 389 37.3%
85–100% 492 31.9%

Llama 3.3 70B Versatile (Groq)

groq-llama-3.3-70b-versatile
35.2%
95% CI 33.6% – 36.8%
  • Sample size3406
  • Correct1198
  • Mean stated confidence 70.4%
  • Confidence gap 35.3 pts
  • Symbols covered 34 / 1064
  • Coverage of the field 3.2%
  • Horizons used3
Calibration — stated confidence against measured accuracy
Confidence n Accuracy
0–50% 25 40.0%
50–70% 718 33.7%
70–85% 2469 35.7%
85–100% 194 33.0%

Over the last 30 days the two confidence intervals do not overlap, so this sample does separate the models.

Accuracy counts a prediction correct when the direction it stated matches the direction the price moved past a 1% threshold over the stated horizon. Coverage is the share of symbols scored in this window that the model expressed an opinion on — a high rate over four symbols is not the same claim as the same rate over ninety. Full method and disclosures