hsb / research

01 / Completed experiments

First, prove
the pipeline.

Synthetic protocol checks and offline ITBench Lite SRE snapshots, run on the IEOR servers at IIT Bombay.

Runs complete

Granite 4.2 3B and the public Param2 17B A2.4B Thinking checkpoint. Complete trajectories were saved for inspection.

Open the experiment repo ↗

Recorded results

One attempt per task in the reported smoke and final incident runs.
TestGranite 4.2 3BParam2 Thinking
Eight synthetic smoke tasks6/8 on the verified rerun5/8
Ordinary tasks, excluding empty-output check5/74/7
Expanded synthetic suite25/32 complete7 direct tasks before interruption
Ten ITBench Lite SRE snapshots0/10 root cause proxy hits0/10 root cause proxy hits
Minimal incident JSON structure2/10 valid10/10 valid
Final incident-run tool calls800

These are focused offline results. The incident score is a deterministic root cause entity proxy using six reference tools. It is not the official ITBench judge score. The interrupted Param2 expanded run does not provide a paired 32-task comparison.

Same score, different failures.

Param2 skipped the evidence. All ten incidents received the same answer, with no tool calls. The answer invented a checkout pod and cited tools it had not used. Structurally valid JSON did not imply a correct diagnosis.

Granite investigated but did not finish reliably. It used all eight investigation turns per incident, made weak or invalid queries, and usually produced malformed final output. The two structurally valid answers still missed the proxy target.

What the smoke run established

The generic model interface, agent loop, tool execution, trajectory logging, and scoring worked end to end. Model protocol reliability remained inconsistent. One deliberately empty-output task counts as a protocol-check pass, not useful task completion.

Reports and evidence

Raw trajectories remain local and on the research server; the committed reports contain run identifiers and result tables.