All projects
No. 6
PackFlow
3D truck loading with reinforcement learning
- My role
- Sole author
- Period
- Jun – Jul 2026
The problem
Loading a delivery truck isn't just fitting boxes: you have to stay under the weight limit, avoid crushing fragile items and keep what gets delivered first near the door. I wanted to see whether a reinforcement learning agent could learn to load better than a simple rule when not everything fits.
What I built
I'm the sole author of the repository: 7 commits, from June 24 to July 14, 2026.
- A voxel Gymnasium environment: a 12 × 6 × 8 truck and a queue of packages. Each action is a position with one of 6 rotations, plus the option to skip the package. The action mask only allows placements that fit, don't collide, rest on at least 80% of their base and stay under the weight limit.
- A reward that adds the placed volume and how well the delivery order is respected, and subtracts damage to fragile packages and the volume left out. The generated packages add up to ~110% to 130% of the truck's capacity, so the agent has to decide what to leave behind.
- Metrics for unloading-order violations and damage, meant for comparing policies.
- Training with MaskablePPO and a custom feature extractor: a 3D CNN for the truck's occupancy and an MLP for the queue. The curriculum has 5 phases, from 4 to 20 packages, and each phase reuses the previous one's weights.
- A FastAPI service with an endpoint that returns the full solution and a WebSocket that streams the loading step by step. If it can't find a PPO checkpoint, it falls back to a greedy baseline that takes the first valid placement.
- A React console built with react-three-fiber that animates the truck with preset scenarios.
- 11 pytest tests for the environment.
Architecture
Results
There are no measured results yet. The repository has no checkpoint and no training logs, so I can't claim the agent beats the greedy baseline. What works today is the environment with its 11 tests, the API with the greedy fallback, and the visualization.
Known limits
- The core step is missing: training the full curriculum and comparing against greedy on loaded volume, delivery order and damage.
- The greedy baseline is weak, since it takes the first valid placement. A fair comparison also needs a stronger heuristic.
- The root README is a copy of the frontend README,
__pycache__is committed and there's a duplicatedpackflow/packflow/apifolder. There is no CI.