Simulation RL Environments

Part 1 of a series on simulation RL environments. AI agents re-plan container-ship dockings at the Port of Barcelona, built from the port's real 2024 records.

Affiliation

Hugging Face

Published

October 6, 2026

PDF

Work in progress

This is an early, experimental series. The environments, numbers and conclusions will change as we go.

Simulation RL environments

Most RL environments for language models are puzzles, code, maths or web tasks. A lot of valuable work looks different. Someone keeps a schedule running under real rules while things keep going wrong. I want agents that can do that work, which raises a general question. How do we build simulation environments, digital twins of real work, that are realistic enough to learn from?

Each part of this series takes one real job, rebuilds it from the records it leaves behind, and grades the agent against an answer that can be proven best.

PortSimEnv v1

PortSimEnv is a deterministic RL environment for berth planning at the Port of Barcelona. It is built from the 1,784 container-ship calls the port recorded in 2024 at its two container terminals, with the real ships, their sizes, berthing times and quay sections. The agent gets one quay, the ships that really called there in a 2024 week, and a week that has just gone wrong. It re-plans the week, deciding for every ship when it docks, where along the quay and with how many cranes.

A recorded eval run, replayed on the 3D twin: GLM-5.3 on week 6 at APM Terminals (17 ships, 9 cranes, two cranes down and two quay closures). It checks three drafts and submits a plan that matches the proven optimum, reward 1.0. The cursor only shows its tool calls. The agent itself sends structured plans.

Every plan plays out on a 3D twin of the real port, built from open data. The agent’s berthing hours, quay sections and crane counts turn into ships, tugs and cranes moving through the week.

The disruptions come from what happens in real ports. Red Sea diversions send ships in late and in bunches, gales stop ships of 300 m or more above 25 knots and every ship above 30, a crane goes down for two days, or an emergency needs a berth now.

Try it

This is the environment’s OpenEnv Space, embedded live. Pick a split and a task index, load the episode and plan it with the same three tools the agent gets. The Playground tab is OpenEnv’s MCP playground. The code is in 07-simulation-environments/portsim-v1 on GitHub.

The PortSimEnv environment, live
It opens on week 7 at APM Terminals (17 ships). Drag ships on the dock chart or type their hour, section and cranes, and the 3D quay follows. Check plan reports rule breaks and cost, 10 times per episode. Submit plan ends the episode with a single grade.

Why it matters

A ship’s containers only move once it is alongside under a crane, so the berth plan decides how long ships wait, how much fuel they burn at anchor and whether a whole service stays on schedule. Schedules rarely hold, so the plan keeps changing.

Barcelona alone handled 3.9 million TEU from 2,163 container-ship calls in 2024, and in 2025 it changed its rules to give berth priority to regular services during congestion. Rotterdam reports that its port-call app cut waiting by 20%.

The eval

Six models each ran the 50 held-out weeks once, with the same three tools, 12 turns and 32k output tokens per turn.

The eval at a glance
Each bar is one model's 50 eval weeks by outcome; the number is its mean reward. Hover for counts.

Watch a rollout

All 300 eval rollouts replay on the 3D quay in the eval Space, FineEnvs/PortSimEnv-Eval, with the plan at each tool call, the dock chart, the grade and the full transcript.

Rollouts on the 3D quay: five eval weeks, six models
Pick a week and a model. The line under the replay explains the reward.

How the reward works

The reward is deterministic, so the same plan always gets the same score. It is computed from the plan alone, without an LLM judge or sampling, once per episode when the agent calls submit_plan.

Reward against cost, with each model's plan
1.0 at the proven optimum, falling towards 0.2 as the plan gets costlier; plans that break a rule sit in the shaded band. Measuring the gap against the avoidable cost means a week full of unavoidable delay is not punished.

An independent checker re-verifies every plan. Across the 1,100 tasks the naive re-plan never scores above 0.44 and a greedy heuristic never above 0.82, because tasks they nearly solve were dropped when the dataset was built.

In the film’s run, GLM-5.3 checks three drafts that come back with 9, 5 and then 0 rule breaks. Its last edits bring NAGOYA EXPRESS down from eight cranes to its six-crane limit and move CMA CGM MONTREAL from hour 131 to 133. The plan it submits costs 87: 72 in weighted delay plus three berth changes at 5 each. That equals the proven optimum, so the gap is zero and the reward is 0.2 + 0.8 × e⁰ = 1.00.

Built on real Port of Barcelona data

Each task keeps a real week at one quay and the ships that really called, adds disruptions modelled on real events, and gets an answer key from CP-SAT. I drop the tasks a simple heuristic nearly solves, which leaves 1,050 training and 50 eval tasks. No week appears in both.

What is real and what is simulated. The records give us the berth plan the port actually ran, with the ships, their sizes, times and quay sections. They don’t include crane assignments, container manifests or time spent waiting at anchor. Workloads, disruptions and the physical handling are simulated, and the film replays an agent’s plan on that simulation.

From seed data to tasks

  1. Seed: the real calls of one quay for one to three weeks, with the berth plan the port actually executed.
  2. Fill the gaps: crane fleets, wind and traffic rules, and a crane rate of 28 moves per hour turn each call’s real time alongside into a workload.
  3. Disturb: inject the disruptions below, scaled by tier (standard, busy, storm, extreme).
  4. Solve: CP-SAT proves the optimal re-plan, the answer key; the naive and greedy re-plans set the floor.
  5. Filter: drop tasks a heuristic nearly solves and keep whole weeks out of training for the eval.
DisruptionWhat happensReal-world modelShare of the 1,100 tasks
Late shipsships arrive hours or days lateRed Sea diversions, half of ships late100%
Closuresquay sections closed for repairs or dredgingcrane-rail repair, fender repair, dredging100%
Crane outagescranes out of service for hours or daysbreakdowns, maintenance89%
Priority cargosome ships count ×3transshipment connections86%
Emergenciesa ship must dock by a deadlinea reefer ship losing power76%
Extra callsunscheduled ships need a berthad-hoc calls63%
Galesno ship moves; ≥300 m stop above 25 kn, all above 30the port’s traffic ordinance62%
Bunchingseveral ships arrive togetherdelayed services catching up46%
Diverted trafficships moved from the other terminalterminal outages15%

Data sources

DataSourceLicence
2024 ship calls: 1,784 at BEST and APM, with times, sections, portsPort of Barcelona open data via berth-allocation-problemsCC BY-SA 4.0
Crane fleets: 13 at BEST, 9 at APMBEST, APM Terminalspublished
Wind, traffic and berth rulesBOE-A-2023-6719, BOE-A-2025-5157, port-call procedurepublic
3D twin: coast, buildings, yards, cranesOpenStreetMapODbL 1.0
TerrainTerrain Tiles on AWSopen
Anchorage and routes (checked)Copernicus Sentinel-2open
Environment: the OpenEnv Space, with the editor and 3D view and an MCP playgroundFineEnvs/PortSimEnv
Eval: all 300 rollouts, replayed in 3DFineEnvs/PortSimEnv-Eval
Dataset: tasks (train, eval), the source calls, the eval rolloutsdatasets/FineEnvs/PortSimEnv
Bucket: 3D twin data and raw eval rolloutsbuckets/FineEnvs/PortSimEnv
CollectionSimulation RL Envs
Codegithub.com/adithya-s-k/FineEnvs: 07-simulation-environments/portsim-v1
DiscussionGitHub discussion #36

Contains data from the Port de Barcelona open data portal. The extracted data and the task pack built from it keep the source’s CC BY-SA 4.0 licence.

Ongoing work

The question behind the series is still open. We can build realistic environments from seed data, but how do we get them as close to the real world as possible? If you have ideas for v2 or v3, for post-training, for other data sources, or a job worth simulating, I’d like to hear them in the GitHub discussion.