Summary
ScienceIDE packages scientific repositories as versioned environments with expert-defined cases, executable tasks and private scientific verifiers. The same environments support supervised fine-tuning, reinforcement learning and evaluation. Verified trajectories train the PhAI-IDE family at 4B, 9B and 72B parameters. The authors report gains on held-out scientific-code repair and selected general coding, reasoning and knowledge benchmarks, and release code plus model artifacts.
Why it matters
For teams building scientific agents, the hard problem is often the environment and verifier rather than the base model. ScienceIDE provides a reusable pattern for converting mature scientific software into executable learning experience. Trial it on one well-tested repository before treating aggregate benchmark gains as evidence of cross-domain scientific ability.
Technical details
| Environment Contract | versioned repository, runtime, expert cases, acceptance criteria and private verifier |
|---|---|
| Training Modes | supervised fine-tuning and reinforcement learning from verified interaction trajectories |
| Models | PhAI-IDE-4B, PhAI-IDE-9B and PhAI-IDE-72B |
| Evaluation | held-out scientific-code repair plus selected public code, reasoning and knowledge benchmarks |
| Artifacts | public code repository and model links |
Tags
ScienceIDEscientific-agentenvironmentverifierreinforcement-learning