hsb / research

Field notes / 2026

Agents, under
inspection.

A working research journal on what models do, where they fail, and the evidence behind every result.

Current investigation

Evaluating the internal Param17B SFT checkpoint. Moving from pipeline validation to a diagnosis of model failures.

Follow the investigation ↗

Research ledger

02 investigations
01
Completed experiments

The pipeline works.
The agents struggle.

Granite 3B and public Param2 Thinking passed basic protocol tasks, then failed ten offline SRE incidents through different mechanisms.

Read the results ↗
  • Location IEOR, IIT Bombay
  • Hardware 2 × RTX A5000
  • ITBench Lite proxy 0/10, both models
  • Period 28 Sep to 01 Oct 2026
02
In preparation

Inside the SFT
failure modes.

Evaluate the internal Param17B checkpoint, inspect the full trajectories, and identify the causes of failure.

See the evaluation plan ↗
  • Location AWS BharatGen HyperPod
  • Cluster bgen-cluster
  • Worker hardware 8 × H100 80 GB
  • New benchmark results Not yet available

Scores are a starting point.

Every conclusion should lead back to an observed answer, a tool call, or a failure in a saved trajectory. The compute ledger records where each experiment ran.