Compare AI models
Two models, one window, one asset class, one horizon. Accuracy is shown with a 95% confidence interval, because a rate without one cannot be compared to another rate.
Rule-Based Analysis
rule-based-analysis
28.9%
95% CI 28.4% – 29.4%
- Sample size32126
- Correct9281
- Mean stated confidence 33.7%
- Confidence gap 4.8 pts
- Symbols covered 1050 / 1064
- Coverage of the field 98.7%
- Horizons used1
| Confidence | n | Accuracy |
|---|---|---|
| 0–50% | 29851 | 28.1% |
| 50–70% | 1394 | 42.2% |
| 70–85% | 389 | 37.3% |
| 85–100% | 492 | 31.9% |
Llama 3.3 70B Versatile (Groq)
groq-llama-3.3-70b-versatile
35.2%
95% CI 33.6% – 36.8%
- Sample size3406
- Correct1198
- Mean stated confidence 70.4%
- Confidence gap 35.3 pts
- Symbols covered 34 / 1064
- Coverage of the field 3.2%
- Horizons used3
| Confidence | n | Accuracy |
|---|---|---|
| 0–50% | 25 | 40.0% |
| 50–70% | 718 | 33.7% |
| 70–85% | 2469 | 35.7% |
| 85–100% | 194 | 33.0% |
Over the last 30 days the two confidence intervals do not overlap, so this sample does separate the models.
Accuracy counts a prediction correct when the direction it stated matches the direction the price moved past a 1% threshold over the stated horizon. Coverage is the share of symbols scored in this window that the model expressed an opinion on — a high rate over four symbols is not the same claim as the same rate over ninety. Full method and disclosures