Compare AI models

Two models, one window, one asset class, one horizon. Accuracy is shown with a 95% confidence interval, because a rate without one cannot be compared to another rate.

Rule-Based Analysis

rule-based-analysis
29.3%
95% CI 28.9% – 29.8%
  • Sample size39532
  • Correct11597
  • Mean stated confidence 33.6%
  • Confidence gap 4.3 pts
  • Symbols covered 1050 / 1064
  • Coverage of the field 98.7%
  • Horizons used1
Calibration — stated confidence against measured accuracy
Confidence n Accuracy
0–50% 36692 28.8%
50–70% 1776 39.2%
70–85% 526 30.2%
85–100% 538 31.2%

Llama 3.3 70B Versatile (Groq)

groq-llama-3.3-70b-versatile
34.9%
95% CI 33.4% – 36.5%
  • Sample size3716
  • Correct1298
  • Mean stated confidence 70.3%
  • Confidence gap 35.3 pts
  • Symbols covered 34 / 1064
  • Coverage of the field 3.2%
  • Horizons used3
Calibration — stated confidence against measured accuracy
Confidence n Accuracy
0–50% 29 41.4%
50–70% 785 33.4%
70–85% 2708 35.5%
85–100% 194 33.0%

Over the last 30 days the two confidence intervals do not overlap, so this sample does separate the models.

Accuracy counts a prediction correct when the direction it stated matches the direction the price moved past a 1% threshold over the stated horizon. Coverage is the share of symbols scored in this window that the model expressed an opinion on — a high rate over four symbols is not the same claim as the same rate over ninety. Full method and disclosures