Compare AI models

Two models, one window, one asset class, one horizon. Accuracy is shown with a 95% confidence interval, because a rate without one cannot be compared to another rate.

Rule-Based Analysis

rule-based-analysis
17.4%
95% CI 16.8% – 18.1%
  • Sample size13769
  • Correct2400
  • Mean stated confidence 32.6%
  • Confidence gap 15.1 pts
  • Symbols covered 1050 / 1064
  • Coverage of the field 98.7%
  • Horizons used1
Calibration — stated confidence against measured accuracy
Confidence n Accuracy
0–50% 13083 16.3%
50–70% 387 44.7%
70–85% 112 44.6%
85–100% 187 25.7%

Llama 3.1 8B Instant (Groq)

groq-llama-3.1-8b-instant
47.2%
95% CI 45.4% – 49.0%
  • Sample size2906
  • Correct1371
  • Mean stated confidence 81.9%
  • Confidence gap 34.7 pts
  • Symbols covered 21 / 1064
  • Coverage of the field 2.0%
  • Horizons used1
Calibration — stated confidence against measured accuracy
Confidence n Accuracy
0–50% 1 0.0%
50–70% 58 10.3%
70–85% 1639 44.1%
85–100% 1208 53.2%

Over the last 30 days the two confidence intervals do not overlap, so this sample does separate the models.

Accuracy counts a prediction correct when the direction it stated matches the direction the price moved past a 1% threshold over the stated horizon. Coverage is the share of symbols scored in this window that the model expressed an opinion on — a high rate over four symbols is not the same claim as the same rate over ninety. Full method and disclosures