🤖 AI benchmark: hit-rate of 7 models

Prematch Live (in-play)

Seven external AI models (Hermes contour) independently analyze the same partite — predicting the esito (1X2), total (Over/Under), both squadre to punteggio (BTTS) and the exact punteggio. Here we honestly compare their pronostici against the real result after the final whistle and combine everything into a single accuracy rating. An informational and analytical snapshot, not betting advice.

⚠️ Data is still accumulating — counting starts from 09.07.2026, so all models are compared on the same events (early test pronostici are excluded). The sample is still small and not representative. Right now the snapshot holds 409 partita(es), 1515 settled AI pronostici (Basketball). The figures below are N, not «a percentage you can trust»: the more partite are played out, the more reliable the snapshot becomes. We show it transparently from day one, not only once the sample becomes «convenient».

Leaderboard · Basketball

Model N (settled) 1X2 Total punti Exact punteggio Composite accuracy
Claude
348 63.5%(221/348) 52.5%(136/259) 0.0%(0/263) 41.0%(357/870)
Kimi
195 63.6%(124/195) 51.5%(84/163) 0.0%(0/163) 39.9%(208/521)
Google AI
438 64.1%(280/437) 47.8%(196/410) 0.0%(0/413) 37.8%(476/1260)
ChatGPT
76 60.5%(46/76) 42.4%(25/59) 0.0%(0/59) 36.6%(71/194)
GLM 5.2
224 63.8%(143/224) 42.5%(94/221) 0.9%(2/222) 35.8%(239/667)
Qwen
197 64.0%(126/197) 35.8%(63/176) 0.6%(1/176) 34.6%(190/549)
DeepSeek
37 51.4%(19/37) 38.2%(13/34) 0.0%(0/37) 29.6%(32/108)

grey — sample <5, not representative; «—» — the model has not made a settled pronostico yet.

Composite accuracy — the share of correct pronostici su all mostrato mercati together: (sum of correct pronostici) ÷ (sum of all settled pronostici) su the mercati 1X2 + Total punti + Exact punteggio. Each mercato-pronostico weighs equally. This is hit-rate, not profitability — for money/ROI by model see /ai-agent. Total: a push (punteggio exactly sulla quota) is excluded from the denominator. «Exact punteggio» — the full final punteggio was guessed correctly (H and A matched); pronostici with no recognized punteggio do not count toward the denominator.

Composite model rating · all mercati · Basketball

Bar height = the model's composite accuracy su all applicable mercati on the current sample. Sorted from best to worst.

41.0% (357/870)
Opus 4.8
39.9% (208/521)
Kimi 2.6
37.8% (476/1260)
Gemini 3.5 Flash
36.6% (71/194)
GPT 5.5
35.8% (239/667)
GLM 5.2
34.6% (190/549)
Qwen 3.7 Plus
29.6% (32/108)
DeepSeek V4 Pro

Bars are AI models by version; grey/dimmed — sample <5, not representative. The snapshot is informational, not betting advice.

Accuracy by mercato · Basketball

Where each model is strong: one mini-bar per applicable mercato, with the percentage and (hits/sample).

Claude Composite 41.0%
1X2
63.5% (221/348)
Total punti
52.5% (136/259)
Exact punteggio
0.0% (0/263)
Kimi Composite 39.9%
1X2
63.6% (124/195)
Total punti
51.5% (84/163)
Exact punteggio
0.0% (0/163)
Google AI Composite 37.8%
1X2
64.1% (280/437)
Total punti
47.8% (196/410)
Exact punteggio
0.0% (0/413)
ChatGPT Composite 36.6%
1X2
60.5% (46/76)
Total punti
42.4% (25/59)
Exact punteggio
0.0% (0/59)
GLM 5.2 Composite 35.8%
1X2
63.8% (143/224)
Total punti
42.5% (94/221)
Exact punteggio
0.9% (2/222)
Qwen Composite 34.6%
1X2
64.0% (126/197)
Total punti
35.8% (63/176)
Exact punteggio
0.6% (1/176)
DeepSeek Composite 29.6%
1X2
51.4% (19/37)
Total punti
38.2% (13/34)
Exact punteggio
0.0% (0/37)

The model's favorite by 1X2 = the max of P1/X/P2 in its probabilities; for sports without a pareggio (tennis, volleyball, etc.) the «X» option doesn't participate. grey — sample <5, not representative. Not betting advice.