Hugging Face has added a dedicated RL Environments filter on the Hugging Face Hub and a set of framework tags to make reinforcement learning tasksets easier to find and run. Environments live as regular dataset repositories: tagging a dataset with rl-environment and one or more framework tags makes it appear in the new filter without creating a new repo type or registry.
Design and scope
The change separates task data from runtimes: tasksets (datasets of tasks, tests, and reward rules) remain on the Hub, while runtime configuration and verifier code can remain in the repository or in external frameworks. The Hub continues to handle hosting, versioning, discovery and previews; frameworks keep providing the execution backends.
Framework tags and running examples
Four frameworks are initially registered as dataset libraries and produce framework-specific snippets on dataset pages: harbor (Harbor), verifiers (Verifiers), openenv (OpenEnv) and nemo-gym (NeMo Gym). A dataset may carry multiple framework tags to indicate compatibility; adding a tag does not change files or convert formats, it only documents which frameworks can run the repo.
Example behaviors described for each framework include:
- Harbor: load task directories from a Hub repo and run a reference solution (the oracle agent) to show verifier output and reward; a Harbor reference used version
harbor==0.21.0in examples. - Verifiers: run the same task directories in different runtimes (for example Docker) and support alternative harnesses such as a minimal bash harness.
- OpenEnv: run a sandboxed agent (example used
openenv[harbor]==0.7.0) and return verifier rewards alongside the agent trace; a reward ofNoneindicates no verifier reward was produced and should be inspected for errors. - NeMo Gym: collect trajectories and compute rewards for evaluation and RL training; example datasets include Structured Outputs, CFBench and SysBench.
To mark a dataset as an environment, add metadata that includes a human-readable name and tags listing rl-environment plus every compatible framework. The team opened PRs to tag several existing repos (examples: Harbor BeyondSWE, Terminal-Lego, Harbor-Mix, NatureBench, Reverse-Text-RL, Multi-SWE-RL-Verified, R2E-Gym-Subset-Verified, Scale-SWE-Verified, NeMo Gym Workplace Assistant, Structured Outputs).
Planned follow-ups include per-config run snippets, structural detection for strict framework layouts and wider automation for tagging during uploads. The Hub’s RL Environments filter is available now at the datasets page with the other=rl-environment parameter for discovery.
Original source: Hugging Face Blog