Compare AI models
Two models, one window, one asset class, one horizon. Accuracy is shown with a 95% confidence interval, because a rate without one cannot be compared to another rate.
Rule-Based Analysis
rule-based-analysis
17.4%
95% CI 16.8% – 18.1%
- Sample size13769
- Correct2400
- Mean stated confidence 32.6%
- Confidence gap 15.1 pts
- Symbols covered 1050 / 1064
- Coverage of the field 98.7%
- Horizons used1
| Confidence | n | Accuracy |
|---|---|---|
| 0–50% | 13083 | 16.3% |
| 50–70% | 387 | 44.7% |
| 70–85% | 112 | 44.6% |
| 85–100% | 187 | 25.7% |
Llama 3.1 8B Instant (Groq)
groq-llama-3.1-8b-instant
47.2%
95% CI 45.4% – 49.0%
- Sample size2906
- Correct1371
- Mean stated confidence 81.9%
- Confidence gap 34.7 pts
- Symbols covered 21 / 1064
- Coverage of the field 2.0%
- Horizons used1
| Confidence | n | Accuracy |
|---|---|---|
| 0–50% | 1 | 0.0% |
| 50–70% | 58 | 10.3% |
| 70–85% | 1639 | 44.1% |
| 85–100% | 1208 | 53.2% |
Over the last 30 days the two confidence intervals do not overlap, so this sample does separate the models.
Accuracy counts a prediction correct when the direction it stated matches the direction the price moved past a 1% threshold over the stated horizon. Coverage is the share of symbols scored in this window that the model expressed an opinion on — a high rate over four symbols is not the same claim as the same rate over ninety. Full method and disclosures