On 11 March 2026, at Cheltenham Festival, our Elite v4 model produced this day at the office:
| Race | Elite's pick | Result | Profit on £10 |
|---|---|---|---|
| 13:20 | King Rasko Grey (11/1) | WON | +£110 |
| 14:00 | Kitzbuhel (11/1) | WON | +£110 |
| 14:40 | Give It To Me Oj (50/1) | 18th | −£10 |
| 15:20 | Final Orders (7/1) | WON | +£70 |
| 16:00 | Quilixios (14/1) | 5th | −£10 |
| 16:40 | Martator (66/1) | WON | +£660 |
| 17:20 | Chicker (80/1) | 18th | −£10 |
4 winners from 7 picks. £920 profit on a £70 outlay. A 1,314% return, on Cheltenham Festival day.
If you saw that, you'd have every reason to think our AI understands Cheltenham. Every reason to want to filter its picks down to a short list of "trusted" courses like this one.
We tested that instinct honestly. It's a trap.
Here's what actually holds up when you test it properly — and what your AI is really doing on every card.
The obvious idea — and why it failed
Our Elite model covers 79 UK, Ireland and France courses. The natural instinct is: filter to the ones where the model performs best. Ignore the tracks where it's losing. Cheltenham looks great — surely we should just back its Cheltenham picks and skip everywhere else?
We tested it. Method:
- Training window — use the last 60 days of Elite v4 predictions to identify the top-performing courses by ROI.
- Test window — measure how those courses actually performed over the next 30 days.
- Rebalance every 30 days, the same way an index fund refreshes its constituents.
- Walk forward across the full model history so we get multiple independent tests.
This is the only honest way to evaluate a filter strategy. You're not allowed to peek at the future when picking your list.
The top-20 course filter, walk-forward tested:
| Test window | Bets | ROI |
|---|---|---|
| Apr 26 → May 25 | 54 | +21.46% |
| May 26 → Jun 24 | 82 | +78.38% |
| Jun 25 → Jul 24 | 305 | −39.56% |
| Jul 25 → Jul 28 | 28 | +39.73% |
| AGGREGATE | 469 | −7.18% |
Two profitable windows. One catastrophic month (−39.56% on 305 bets — the dominant sample). Net loss.
The tighter refinement — top 10 courses
If top-20 diluted the signal with weaker picks, surely top-10 concentrates it? We ran the same walk-forward test with a top-10 filter.
The top-10 course filter, walk-forward tested:
| Test window | Bets | ROI |
|---|---|---|
| Apr 26 → May 25 | 15 | −3.33% |
| May 26 → Jun 24 | 39 | +17.16% |
| Jun 25 → Jul 24 | 165 | −46.94% |
| Jul 25 → Jul 28 | 15 | +0.83% |
| AGGREGATE | 234 | −30.40% |
Worse. The tighter filter didn't rescue the strategy — it amplified the drawdown. −30% ROI over 234 bets is what backing a losing tipster looks like.
At this point, most content-hungry AI racing sites would quietly file this away and publish a different article about how great their "top 10 courses" system is. We're publishing it because it matters.
The counter-intuitive finding
Here's what actually delivered profit over the same period — backing every single rank-1 pick, no filter at all:
| Test window | Bets | Strike | Elite ROI | SP Favourite ROI | Edge |
|---|---|---|---|---|---|
| Feb 25 → Mar 26 | 882 | 21.4% | +36.77% | −8.60% | +45.4pp |
| Mar 27 → Apr 25 | 672 | 27.8% | +53.31% | −8.79% | +62.1pp |
| Apr 26 → May 25 | 403 | 21.8% | −4.12% | −9.95% | +5.8pp |
| May 26 → Jun 24 | 415 | 21.0% | +29.44% | −11.41% | +40.9pp |
| Jun 25 → Jul 24 | 953 | 15.6% | −20.32% | −12.50% | −7.8pp |
| Jul 25 → Jul 28 | 105 | 13.3% | −10.15% | −13.96% | +3.8pp |
| AGGREGATE | 3,430 | 20.8% | +17.02% | −10.51% | +27.5pp |
No filter. Just back every rank-1 pick. In this 155-day walk-forward test:
- 3,430 walk-forward bets — substantially larger than the course-level samples we tested earlier in this article
- +17.02% ROI — a level-stakes punter putting £10 on every pick would be +£5,838 up over the period
- 20.8% strike rate at an average price around 7/1
- +27.5 percentage-point edge over blindly backing the SP favourite (a bettor's cheapest possible baseline)
- 5 of 6 walk-forward months beat the market baseline
- 4 of 6 were flat-out profitable
What that looks like on the equity curve
The dashed red line is what a punter blindly backing the SP favourite would have done: a slow bleed to −£6,395. The green line is Elite v4's rank-1 picks over the same days — trending up despite one obvious drawdown in late June/early July. A £12,232 swing between the two strategies over 141 trading days.
The one bad month (Jun 25 → Jul 24) lost 20% ROI. It happens. Every real signal has drawdowns. What matters is whether the strategy recovers — and it did.
An important caveat before you get excited
3,430 bets over 155 days is a much more reliable sample than the course-level slices we looked at earlier, but it isn't proof of a permanent edge. Markets evolve, models decay, and future performance may differ from what we've measured here. What this test shows is that the signal survived a large out-of-sample walk-forward evaluation and substantially outperformed a simple market benchmark over the period tested. That's the honest claim.
The edges most likely to survive going forward are structural ones — features of the model's underlying inputs (form, class movement, going, connections) that reflect genuine patterns in the sport — rather than recent performance streaks on specific tracks. That's what course-filtering was: a bet on recent performance streaks. That's why it failed.
Why filtering broke what worked
Small samples are mostly noise. A course with 15–30 bets can look like a "strong edge" or a "heavy loss" depending on which 5 races land differently. When we filter to "top 10 courses," we're picking the 10 courses that got lucky in the training window, not the 10 where the model has genuine skill. And courses that got lucky tend to regress.
The full model is where the skill lives. Elite v4's edge comes from its underlying probability estimates — how it weighs form, class movement, going preferences, trainer form, jockey booking, pedigree signals. Those signals apply across every course, in aggregate. Slicing the model into per-course subsets doesn't make it smarter. It just makes it noisier.
The parallel with investing is imperfect but instructive: a diversified index tends to beat a "top 10 stocks by last quarter's return" portfolio because recent winners regress and the underlying market average reasserts itself.
What this means for how you use RaceChatAI
Course-based cherry-picking underperformed. The temptation to say "our AI is best at Cheltenham" or "we should skip Wolverhampton" is real — but the walk-forward evidence points the other way.
The historical evidence suggests the model performs best when evaluated across all rank-1 selections, rather than after applying course-based filters. When RaceChatAI surfaces a rank-1 pick, back it or don't based on the price, not on which course it's at.
Respect variance. +17% ROI over 155 days doesn't mean +17% every week. Some weeks the picks fire; some weeks they don't. The edge is aggregate, over many bets, across a full season. Level-stake or fractional-Kelly it; don't chase or fade it based on the last three results. Past performance is not a promise of future performance.
Try it in the chat
Ask RaceChatAI:
"What's Elite's rank-1 pick in every race today?"
You'll get the full list — every course, every race, every rank-1 the model rates. That's the strategy the historical evidence supports. Not a curated subset. The whole card.
The honest bottom line
Three strategies tested. Two failed:
- Top-20 course filter — lost 7% out-of-sample.
- Top-10 course filter — lost 30% out-of-sample.
One retained a positive edge in this walk-forward evaluation:
- Back every rank-1 pick — +17.02% ROI over 3,430 walk-forward bets across 155 days, +27.5pp better than the market baseline.
That's the process. Most AI racing sites won't publish the failures. That's how their "systems" tend to be Cheltenham at +1,314% ROI in bold letters — until you check the base rate and realise you're being sold noise.
Our promise isn't that we'll pick every winner. It's that when we say something works, we've tested it in a way that would embarrass us if it didn't.
In our walk-forward testing, the full unfiltered rank-1 strategy was the only approach that consistently retained its edge out of sample.
Methodology — how the numbers were produced
For readers who want to reproduce or stress-test the claim:
- Data source. All picks and results come from our internal
elite_predictionsandhorsestables. Each row is the model's pre-race rank-1 selection plus the horse's actual finishing position and settled Starting Price. Prices are populated from the Racing API's post-race payload — the industry SP, not Betfair SP or early morning prices. Every ROI figure would be somewhat higher on BSP after Betfair commission. - Bet universe. Every Elite v4 rank-1 pick over the 155 day window (25 Feb 2026 → 28 Jul 2026) where the horse had a settled finishing position AND a parseable fractional SP. Rows without an SP or listed as
SPunpriced are excluded. Non-runners (position IS NULL) are naturally excluded. - Scope. UK, Ireland and France racing only. Hong Kong, USA and Australia meetings are filtered out because the model was not trained on them and we don't serve predictions there.
- Model lineage. Elite v4 is retrained periodically — v6 through v14 were live for various stretches of the tested window. This is real live-serving performance (not one static model backtested), which is stronger than a frozen backtest: every prediction was made in real time by whichever version was live that day. The full version history and per-version validation logloss is visible in the admin panel.
- Walk-forward split. The article groups the 155 days into six ~30-day test windows. No look-back optimisation was applied — the aggregate is just the concatenation of each window's real-time picks and results.
- ROI formula. Level £10 stakes. Win →
+£10 × (decimal_odds − 1). Loss →−£10. Rounded to whole pounds for display. ROI = total profit ÷ total staked × 100. - SP Favourite baseline. For each race in the same window, we back the horse with the lowest settled SP (tie-break: same query as picks). Same odds-parsing, same £10 stakes.
- Chart. The equity curve at the top is the cumulative running total of £10-per-bet profit for each of the two strategies, plotted per race-day.
- Known limitations. Real-world execution would face bet limits from bookmakers on winning selections, and industry SP may be shorter than the price you actually got on the day. Live-forward performance since the article was published is tracked separately and rebuilt nightly at 05:00 UTC — ask "how has your AI performed?" in the chat to see the current numbers.
If any of the above raises a question, get in touch — we'd rather answer a sceptic than have them silently distrust the numbers.
RaceChatAI — the AI horse racing analyst that shows its working. Chat live at racechatai.com.
This article describes historical model performance in a walk-forward evaluation. Racing is variance-heavy and no strategy guarantees future returns. Only bet what you can afford to lose.