Summary

ScienceIDE packages scientific repositories as versioned environments with expert-defined cases, executable tasks and private scientific verifiers. The same environments support supervised fine-tuning, reinforcement learning and evaluation. Verified trajectories train the PhAI-IDE family at 4B, 9B and 72B parameters. The authors report gains on held-out scientific-code repair and selected general coding, reasoning and knowledge benchmarks, and release code plus model artifacts.

Why it matters
For teams building scientific agents, the hard problem is often the environment and verifier rather than the base model. ScienceIDE provides a reusable pattern for converting mature scientific software into executable learning experience. Trial it on one well-tested repository before treating aggregate benchmark gains as evidence of cross-domain scientific ability.
Technical details
Environment Contract versioned repository, runtime, expert cases, acceptance criteria and private verifier
Training Modes supervised fine-tuning and reinforcement learning from verified interaction trajectories
Models PhAI-IDE-4B, PhAI-IDE-9B and PhAI-IDE-72B
Evaluation held-out scientific-code repair plus selected public code, reasoning and knowledge benchmarks
Artifacts public code repository and model links
Tags
ScienceIDEscientific-agentenvironmentverifierreinforcement-learning