Analysis report — bare-gemma4-s11-20260721
Run report — bare-gemma4-s11-20260721
Entrant: bare-gemma4 · class: local-20gb · kind: course · passed: NO · started: 2026-07-22T06:31:42
Coach's read
The agent failed to complete the run and did not pass the obstacle course. During room one, the agent spent almost the entire split time between tool calls rather than within them, suggesting a need to tighten the loop. The referee recorded 11 faults, most notably making expensive submissions without providing evidence. To improve, you should replay the ledger to find where the path was lost and focus on clearing the failed room.
Splits
| room | split (s) | tool time (s) | calls | faults | field median | field best |
|---|---|---|---|---|---|---|
| room1 | 41.92 | 0.01 | 39 | 11 | 65.287 | 46.94 |
Field: 2 comparable passed course run(s).
What cost time
- room1: 19 duplicate call(s); top stall 2.8s (run_sql → submit_finding)
Three things to fix
- The run did not finish — everything else is secondary to clearing the failed room. Replay the ledger and find where the path was lost.
- room1: 41.9s of the 41.9s split was spent between tool calls (thinking), not in them — tighten the loop.
- room1: 11 fault(s) recorded by the referee. Submissions without evidence are the most expensive mistake on the course.
Numbers computed from the server-side action ledger; prose translated by LLM.