Gauntlet test report — anomaly review

Stockfish 17, in the dark

A dominant 30-engine gauntlet, and two losses that don't look like chess mistakes — they look like a clock running out.

69.7%
overall score / 1500 games
+9.4
live Elo vs. start
2
losses under review
9+0.1
time control (5Pct of 3+2 blitz)

Depth, not eval, is where this breaks

In both losses, Stockfish's own reported search depth collapses to 1–6 ply for extended stretches — exactly where the position falls apart. The evaluation stays confidently positive throughout; it's simply not backed by real search.

GAME 1 — vs PlentyChess 2.0.0 (start 2801) · White search depth, moves 20–64 0–1, adjudicated move 64
move 20move 64
normal depth (≥10 ply) collapse (<10 ply)
44. Bf5 {+2.78/15} 45. Bd2 {+2.78/6} 46. Bxe3 {+3.65/2} 47. Rfg1 {+2.43/5} 48. Qg4 {+1.51/6} 49. Rg3 {+0.33/1} 50. Rf1 {-1.31/4} 51. Kg2 {-6.93/20} ← depth recovers, already lost
In that six-move window, Black's connected e/f-pawns run down the board and land 50...e2! 51...Rxh2+! 52...exf1=N+! — an underpromotion fork winning material outright. Stockfish never reached the depth needed to see it coming.

A 112-move game, half of it at depth 1

Against RubiChess 20221203 — rated 171 points below Stockfish — the same signature appears, but sustained far longer: from roughly move 47 onward, across most of the remaining game, White is repeatedly reduced to a single ply of search.

GAME 2 — vs RubiChess 20221203 (start 2709) · White search depth, moves 6–112 0–1, adjudicated move 113
move 6move 112
normal depth (≥10 ply) collapse (<10 ply)
The eval sits comfortably around +1 to +2 for nearly the entire second half of the game — Stockfish genuinely was better — but with no search behind it, it can't convert. Only in the final three moves does the position actually give way: eval drops from +17.85 to -199.95, and the game ends by adjudication.

Likely cause and context

Working theory

9+0.1 — the "5Pct" tag, 5% of a standard 3+2 blitz baseline — gives a full-game budget of only about 9 + (moves × 0.1) seconds. For Game 2's 112 moves that's ~20 seconds total, for the entire game. This rig (i5-4570) is also noticeably older than other hardware tested. Both games show deep, healthy search (15–23 ply) early on, followed by a sudden drop to near-zero — consistent with overspending that tiny budget early and running dry for the rest of the game, rather than a genuine evaluation or strength failure.

Not representative

Stockfish still won the 1,500-game event comfortably — 69.7% score, a wide margin over 2nd place. These two losses read as rare time-trouble outliers, not a systemic weakness. Worth a look if depth-1 stretches like these show up in time-allocation logs for this TC/hardware pairing.