hsb / research

Infrastructure / Experiment provenance

Where the
evidence runs.

A compute ledger connecting each experiment to its hardware, location, and execution status.

Two environments

The first pipeline tests ran on institute hardware. The next SFT evaluation is being prepared on BharatGen H100 compute.

Browse completed runs ↗
01
Completed experiments

IEOR, IIT Bombay

passpoli research server

Hardware2 NVIDIA RTX A5000 GPUs, 24 GB VRAM each
Models runGranite 4.2 3B and public Param2 17B A2.4B Thinking
ExperimentsSynthetic smoke tests, expanded protocol checks, and ten ITBench Lite SRE snapshots
ServingGranite through Ollama; Param2 through Transformers
StatusReported experiments complete; interrupted runs labeled separately
Results and failure analysis ↗
02
Current preparation

BharatGen on AWS

SageMaker HyperPod, Mumbai region

Clusterbgen-cluster
Worker typeml.p5.48xlarge: 8 NVIDIA H100 GPUs, 80 GB each
Model targetInternal Param17B step_20000 SFT checkpoint
Storage and runtimeFSx project storage; Docker experiments on compute workers
VerifiedAn idle worker candidate and cached-container startup
PendingModel runtime, native tool protocol, benchmark runs, and scores
Current investigation ↗

Environment boundaries

The institute smoke environment used an isolated Python environment. The AWS workflow uses Docker on compute workers with project data on FSx. The login node is a gateway. The earlier server scripts require adaptation before use on AWS.