This is an early, experimental series. The environments, numbers and conclusions will change as we go.
Simulation RL environments
Most RL environments for language models are puzzles, code, maths or web tasks. A lot of valuable work looks different. Someone keeps a schedule running under real rules while things keep going wrong. I want agents that can do that work, which raises a general question. How do we build simulation environments, digital twins of real work, that are realistic enough to learn from?
Each part of this series takes one real job, rebuilds it from the records it leaves behind, and grades the agent against an answer that can be proven best.
PortSimEnv v1
PortSimEnv is a deterministic RL environment for berth planning at the Port of Barcelona. It is built from the 1,784 container-ship calls the port recorded in 2024 at its two container terminals, with the real ships, their sizes, berthing times and quay sections. The agent gets one quay, the ships that really called there in a 2024 week, and a week that has just gone wrong. It re-plans the week, deciding for every ship when it docks, where along the quay and with how many cranes.
Every plan plays out on a 3D twin of the real port, built from open data. The agent’s berthing hours, quay sections and crane counts turn into ships, tugs and cranes moving through the week.
The disruptions come from what happens in real ports. Red Sea diversions send ships in late and in bunches, gales stop ships of 300 m or more above 25 knots and every ship above 30, a crane goes down for two days, or an emergency needs a berth now.
Try it
This is the environment’s OpenEnv Space, embedded live. Pick a split and a task index, load the episode and plan it with the same three tools the agent gets. The Playground tab is OpenEnv’s MCP playground. The code is in 07-simulation-environments/portsim-v1 on GitHub.
Why it matters
A ship’s containers only move once it is alongside under a crane, so the berth plan decides how long ships wait, how much fuel they burn at anchor and whether a whole service stays on schedule. Schedules rarely hold, so the plan keeps changing.
Barcelona alone handled 3.9 million TEU from 2,163 container-ship calls in 2024, and in 2025 it changed its rules to give berth priority to regular services during congestion. Rotterdam reports that its port-call app cut waiting by 20%.
The eval
Six models each ran the 50 held-out weeks once, with the same three tools, 12 turns and 32k output tokens per turn.
- GPT-6.1 Sol leads at 0.89 and submits a valid plan every week, but it matches the optimum on only 26 of 50, so there is plenty of headroom.
- Every model scores lower on storm and extreme weeks, which stack gales, closures and emergencies.
- The open models lose most of their score before grading. They spend their output budget reasoning and never call
submit_plan. On the 13 to 25 weeks where they do submit a valid plan, they average 0.75 to 0.97, so getting them to submit is the first thing to train.
Watch a rollout
All 300 eval rollouts replay on the 3D quay in the eval Space, FineEnvs/PortSimEnv-Eval, with the plan at each tool call, the dock chart, the grade and the full transcript.
How the reward works
The reward is deterministic, so the same plan always gets the same score. It is computed from the plan alone, without
an LLM judge or sampling, once per episode when the agent calls submit_plan.
An independent checker re-verifies every plan. Across the 1,100 tasks the naive re-plan never scores above 0.44 and a greedy heuristic never above 0.82, because tasks they nearly solve were dropped when the dataset was built.
In the film’s run, GLM-5.3 checks three drafts that come back with 9, 5 and then 0 rule breaks. Its last edits bring
NAGOYA EXPRESS down from eight cranes to its six-crane limit and move CMA CGM MONTREAL from hour 131 to 133. The plan
it submits costs 87: 72 in weighted delay plus three berth changes at 5 each. That equals the proven optimum, so
the gap is zero and the reward is 0.2 + 0.8 × e⁰ = 1.00.
Built on real Port of Barcelona data
Each task keeps a real week at one quay and the ships that really called, adds disruptions modelled on real events, and gets an answer key from CP-SAT. I drop the tasks a simple heuristic nearly solves, which leaves 1,050 training and 50 eval tasks. No week appears in both.
What is real and what is simulated. The records give us the berth plan the port actually ran, with the ships, their sizes, times and quay sections. They don’t include crane assignments, container manifests or time spent waiting at anchor. Workloads, disruptions and the physical handling are simulated, and the film replays an agent’s plan on that simulation.
From seed data to tasks
- Seed: the real calls of one quay for one to three weeks, with the berth plan the port actually executed.
- Fill the gaps: crane fleets, wind and traffic rules, and a crane rate of 28 moves per hour turn each call’s real time alongside into a workload.
- Disturb: inject the disruptions below, scaled by tier (standard, busy, storm, extreme).
- Solve: CP-SAT proves the optimal re-plan, the answer key; the naive and greedy re-plans set the floor.
- Filter: drop tasks a heuristic nearly solves and keep whole weeks out of training for the eval.
| Disruption | What happens | Real-world model | Share of the 1,100 tasks |
|---|---|---|---|
| Late ships | ships arrive hours or days late | Red Sea diversions, half of ships late | 100% |
| Closures | quay sections closed for repairs or dredging | crane-rail repair, fender repair, dredging | 100% |
| Crane outages | cranes out of service for hours or days | breakdowns, maintenance | 89% |
| Priority cargo | some ships count ×3 | transshipment connections | 86% |
| Emergencies | a ship must dock by a deadline | a reefer ship losing power | 76% |
| Extra calls | unscheduled ships need a berth | ad-hoc calls | 63% |
| Gales | no ship moves; ≥300 m stop above 25 kn, all above 30 | the port’s traffic ordinance | 62% |
| Bunching | several ships arrive together | delayed services catching up | 46% |
| Diverted traffic | ships moved from the other terminal | terminal outages | 15% |
Data sources
| Data | Source | Licence |
|---|---|---|
| 2024 ship calls: 1,784 at BEST and APM, with times, sections, ports | Port of Barcelona open data via berth-allocation-problems | CC BY-SA 4.0 |
| Crane fleets: 13 at BEST, 9 at APM | BEST, APM Terminals | published |
| Wind, traffic and berth rules | BOE-A-2023-6719, BOE-A-2025-5157, port-call procedure | public |
| 3D twin: coast, buildings, yards, cranes | OpenStreetMap | ODbL 1.0 |
| Terrain | Terrain Tiles on AWS | open |
| Anchorage and routes (checked) | Copernicus Sentinel-2 | open |
Links
| Environment: the OpenEnv Space, with the editor and 3D view and an MCP playground | FineEnvs/PortSimEnv |
| Eval: all 300 rollouts, replayed in 3D | FineEnvs/PortSimEnv-Eval |
| Dataset: tasks (train, eval), the source calls, the eval rollouts | datasets/FineEnvs/PortSimEnv |
| Bucket: 3D twin data and raw eval rollouts | buckets/FineEnvs/PortSimEnv |
| Collection | Simulation RL Envs |
| Code | github.com/adithya-s-k/FineEnvs: 07-simulation-environments/portsim-v1 |
| Discussion | GitHub discussion #36 |
Contains data from the Port de Barcelona open data portal. The extracted data and the task pack built from it keep the source’s CC BY-SA 4.0 licence.
Ongoing work
The question behind the series is still open. We can build realistic environments from seed data, but how do we get them as close to the real world as possible? If you have ideas for v2 or v3, for post-training, for other data sources, or a job worth simulating, I’d like to hear them in the GitHub discussion.