Compare AI models

Two models, one window, one asset class, one horizon. Accuracy is shown with a 95% confidence interval, because a rate without one cannot be compared to another rate.

Rule-Based Analysis

rule-based-analysis
29.5%
95% CI 29.1% – 30.0%
  • Sample size44450
  • Correct13133
  • Mean stated confidence 33.6%
  • Confidence gap 4.0 pts
  • Symbols covered 1052 / 1066
  • Coverage of the field 98.7%
  • Horizons used1
Calibration — stated confidence against measured accuracy
Confidence n Accuracy
0–50% 41343 29.0%
50–70% 1877 39.5%
70–85% 601 30.1%
85–100% 629 36.7%

Llama 3.3 70B Versatile (Groq)

groq-llama-3.3-70b-versatile
35.0%
95% CI 33.7% – 36.4%
  • Sample size4573
  • Correct1602
  • Mean stated confidence 70.0%
  • Confidence gap 35.0 pts
  • Symbols covered 34 / 1066
  • Coverage of the field 3.2%
  • Horizons used3
Calibration — stated confidence against measured accuracy
Confidence n Accuracy
0–50% 34 44.1%
50–70% 1010 32.9%
70–85% 3332 35.7%
85–100% 197 33.5%

Over the last 30 days the two confidence intervals do not overlap, so this sample does separate the models.

Accuracy counts a prediction correct when the direction it stated matches the direction the price moved past a 1% threshold over the stated horizon. Coverage is the share of symbols scored in this window that the model expressed an opinion on — a high rate over four symbols is not the same claim as the same rate over ninety. Full method and disclosures