Recorded results
| Test | Granite 4.2 3B | Param2 Thinking |
|---|---|---|
| Eight synthetic smoke tasks | 6/8 on the verified rerun | 5/8 |
| Ordinary tasks, excluding empty-output check | 5/7 | 4/7 |
| Expanded synthetic suite | 25/32 complete | 7 direct tasks before interruption |
| Ten ITBench Lite SRE snapshots | 0/10 root cause proxy hits | 0/10 root cause proxy hits |
| Minimal incident JSON structure | 2/10 valid | 10/10 valid |
| Final incident-run tool calls | 80 | 0 |
These are focused offline results. The incident score is a deterministic root cause entity proxy using six reference tools. It is not the official ITBench judge score. The interrupted Param2 expanded run does not provide a paired 32-task comparison.
Same score, different failures.
Param2 skipped the evidence. All ten incidents received the same answer, with no tool calls. The answer invented a checkout pod and cited tools it had not used. Structurally valid JSON did not imply a correct diagnosis.
Granite investigated but did not finish reliably. It used all eight investigation turns per incident, made weak or invalid queries, and usually produced malformed final output. The two structurally valid answers still missed the proxy target.
What the smoke run established
The generic model interface, agent loop, tool execution, trajectory logging, and scoring worked end to end. Model protocol reliability remained inconsistent. One deliberately empty-output task counts as a protocol-check pass, not useful task completion.
Reports and evidence
Raw trajectories remain local and on the research server; the committed reports contain run identifiers and result tables.