Field notes / 2026
Agents, under
inspection.
A working research journal on what models do, where they fail, and the evidence behind every result.
Current investigationEvaluating the internal Param17B SFT checkpoint. Moving from pipeline validation to a diagnosis of model failures.
Follow the investigation ↗ Research ledger
02 investigations
01
Completed experimentsGranite 3B and public Param2 Thinking passed basic protocol tasks, then failed ten offline SRE incidents through different mechanisms.
Read the results ↗ - Location IEOR, IIT Bombay
- Hardware 2 × RTX A5000
- ITBench Lite proxy 0/10, both models
- Period 28 Sep to 01 Oct 2026
02
In preparationEvaluate the internal Param17B checkpoint, inspect the full trajectories, and identify the causes of failure.
See the evaluation plan ↗ - Location AWS BharatGen HyperPod
- Cluster bgen-cluster
- Worker hardware 8 × H100 80 GB
- New benchmark results Not yet available
Scores are a starting point.
Every conclusion should lead back to an observed answer, a tool call, or a failure in a saved trajectory. The compute ledger records where each experiment ran.