Compare AI models
Two models, one window, one asset class, one horizon. Accuracy is shown with a 95% confidence interval, because a rate without one cannot be compared to another rate.
Rule-Based Analysis
rule-based-analysis
29.5%
95% CI 29.1% – 30.0%
- Sample size44450
- Correct13133
- Mean stated confidence 33.6%
- Confidence gap 4.0 pts
- Symbols covered 1052 / 1066
- Coverage of the field 98.7%
- Horizons used1
| Confidence | n | Accuracy |
|---|---|---|
| 0–50% | 41343 | 29.0% |
| 50–70% | 1877 | 39.5% |
| 70–85% | 601 | 30.1% |
| 85–100% | 629 | 36.7% |
Llama 3.3 70B Versatile (Groq)
groq-llama-3.3-70b-versatile
35.0%
95% CI 33.7% – 36.4%
- Sample size4573
- Correct1602
- Mean stated confidence 70.0%
- Confidence gap 35.0 pts
- Symbols covered 34 / 1066
- Coverage of the field 3.2%
- Horizons used3
| Confidence | n | Accuracy |
|---|---|---|
| 0–50% | 34 | 44.1% |
| 50–70% | 1010 | 32.9% |
| 70–85% | 3332 | 35.7% |
| 85–100% | 197 | 33.5% |
Over the last 30 days the two confidence intervals do not overlap, so this sample does separate the models.
Accuracy counts a prediction correct when the direction it stated matches the direction the price moved past a 1% threshold over the stated horizon. Coverage is the share of symbols scored in this window that the model expressed an opinion on — a high rate over four symbols is not the same claim as the same rate over ninety. Full method and disclosures