Growing Training Grounds from Code Itself: How CodeMidas Turns Implemented Functionality into RL Environments for Coding Agents
Synopsis
CodeMidas presents an agentic pipeline that uses source code as its only task-specific input to turn implemented functionality in open-source codebases into executable, verifiable coding RL environments, yielding 5,545 training tasks from 3,185 codebases across 23 programming languages and 15 technical domains, and training MiMo-V2.5 with GRPO improves all five external benchmarks, including DeepSWE pass rate from 10.0% to 21.7%, ProgramBench Almost Solved from 4.5 to 21.5, and Terminal-Bench v2.1 from 63.7% to 72.2%.
Interpretation
Treating implemented functionality itself as the task source: agents inspect codebase structure and build metadata to identify functionality with public entry points and observable outcomes, trace public entry points and shared dependencies to define task scope, remove the selected core implementation, and adjust the remaining code into a coherent starting point while retaining the original implementation as a reference solution. Existing pipelines largely derive tasks from development artifacts (issues, pull requests, commits), existing tests, or documentation, tying task coverage to the coverage of those records; CodeMidas uses source code as its only task-specific input, extending the extractable range of tasks to implemented functionality that lacks such records. The paper compares task-specific input requirements item by item in Table 1; CodeMidas satisfies all five conditions (no issue, PR, commit, existing tests, or written description) and covers 23 languages, the broadest in that comparison.
Tests are anchored in execution of the original code: agents map the statement's behavioral requirements to test inputs and boundary cases, invoke public entry points in a reference copy and record outcomes, using command executions for CLI tools, input-output cases for pure functions, and sequences of calls for stateful APIs; every assertion is then reviewed to remove restrictions unsupported by the statement, such as exact wording, incidental ordering, or internal structure, and a task is rejected if an assertion depends on a private symbol with no behavioral substitute. This lets the verifier reject incorrect solutions while still accepting alternative correct implementations, directly addressing the test-oracle problem; the paper cites EvalPlus and PatchDiff to note that expanded tests uncover incorrect generated programs missed by original suites and that incorrect patches have been accepted by SWE-bench tests. The method specifies assertion review and rerunning revised tests on the reference solution; at the environment level, each task is checked in six fresh containers for execution consistency, with both starting-state runs required to fail and all four reference runs required to pass.
Environments are filtered with agent rollouts before training: adversarial rollouts search the solver-visible environment, including compiled artifacts, caches, files left by construction agents, and installed copies of the target project, logging commands and outputs for each suspected exploit for separate review; a coding agent attempts each task four times and a reviewing agent judges implementations against the statement, verifier, and reference solution, flagging false positives and false negatives; tasks are then kept only when rollouts show both successful and failed attempts. This moves environment reliability beyond static execution checks toward rollout-based leakage probing and verifier-agreement review, and explicitly notes that all-pass or all-fail outcomes do not reveal their cause, which may be difficulty or remaining defects such as weak tests or requirements missing from the statement, so they are not used as a retention criterion. The three filtering steps are described individually in the method; ablations show the high-quality 5k pool exceeds the roughly 8k vanilla sample by 0.59, 4.59, and 4.49 percentage points on SWE-bench Pro, DeepSWE, and CodeMidas Val, and even the high-quality 3k subset outperforms that 8k sample on all three evaluations.
Training and behavior analysis: MiMo-V2.5 is trained with GRPO, binary execution rewards, batch size 32, and 32 rollouts per task; trajectory analysis shows agents explore more (pre-edit read/search calls rising from 27.2 to 40.1), show greater overlap between written code and preceding reasoning (drafting ratio from 0.358 to 0.629), and run more distinct verification commands after the final edit (2.03 to 2.53), with agent-written and executed checks associated with a mean pass rate 4.2 percentage points higher within the same task and checkpoint (95% CI: 1.8-6.6). It links benchmark gains to observable behavioral changes and compares early and late checkpoints on SWE-bench Pro, ProgramBench, and Terminal-Bench v2.1, showing that changes such as increased exploration also appear on external tasks. Behavioral metrics are defined in Appendix B (exploration counted as deduplicated read/search requests, drafting by sampling 16-character fragments at a stride of four, self-verification as distinct verification commands after the final edit); confidence intervals resample whole tasks, keeping all observed checkpoints from each sampled task together.
Perspective
This work targets training and data-engineering settings where RL environments must be built for coding agents: it applies to open-source codebases with executable build processes, public entry points, and observable behavior, covering command-line tools, pure library functions, and stateful library APIs, with evaluation spanning issue repair, whole-program construction, code translation, and terminal work. It lets follow-up work expand task pools without relying on issues, pull requests, commits, existing tests, or written descriptions, and treats environment reliability as an actionable filtering target; for readers focused only on inference-time prompting or non-coding tasks, direct transferability is limited.
A careful reader would still watch: how filtered tasks behave on new codebases outside the training distribution; how much of the discarded all-pass or all-fail tasks reflects difficulty versus defects, which the paper explicitly says these outcomes cannot reveal; whether the association between self-verification and higher pass rates, established within the same task and checkpoint, has a causal direction that still needs characterization; and that the loaded text is the full paper text, so figures such as language and domain distributions, reference-solution size bins, and learning curves appear as prose descriptions rather than directly visible graphics.
