hsb / research

02 / Current investigation

Find the failure.
Explain the cause.

Evaluate the internal Param17B SFT checkpoint and inspect why its agent behavior goes wrong.

Preparation verified

The worker and Docker startup are verified. Model serving, native tool calling, and new benchmark results are still pending.

Open the evaluation repo ↗

The research question

When the SFT model fails a task, what caused it? Distinguish serving and parsing defects from skipped evidence, incorrect tool use, reasoning errors, and malformed final answers. Manual inspection of every result supports that diagnosis.

Evaluation sequence

  1. Validate the runtimeEstablish a compatible model load and preserve the checkpoint, image, tokenizer, and generation configuration.
  2. Verify the native protocolInspect a direct answer, a tool call, and a chained tool task before committing to a longer evaluation.
  3. Run ITBenchFreeze case selection and budgets. Separate proxy scoring from any official judge evaluation.
  4. Add complementary testsSelect two or three additional tests covering tool use and reasoning. GDPval is a candidate; its protocol and grading route are not finalized.
  5. Inspect every outcomeKeep raw outputs and artifacts, record manual verdicts separately from automatic scores, and categorize failures.

No new scores yet. Capacity and Docker checks are infrastructure evidence. They are not a completed model evaluation or a resource reservation.

Expected research artifacts

A findings report, benchmark manifests, raw trajectories, generated task artifacts, a local review viewer, manual verdicts, and exact reproduction commands. Any intervention that forces evidence consumption or tool calls will be reported separately from the baseline.