René Cano
All projects

No. 11st place · Expo Ingenierías 2026

UrbanMind X

Traffic light control with reinforcement learning

My role
Lead author (a teammate built the SUMO network)
Period
May – Jun 2026
UrbanMind X dashboard cover: the “UrbanMind·X” title over a 3D city of lit blocks, with the tagline “A city that learns to clear itself”.

The problem

Fixed-time traffic lights repeat the same cycle even when one avenue is jammed and the other is empty. I wanted to test whether an agent that watches traffic in real time makes better decisions about when to switch phases, at a real Toluca intersection: Venustiano Carranza and Blvd. Pino Suárez, using traffic counts from the Quivera study (UAEM, 2022).

What I built

I'm the main author of the repository: 28 commits and 4,911 lines added. A teammate built the SUMO network and routes (4 commits, 1,137 lines).

  • A TraCI wrapper that connects Python to the SUMO simulation.
  • Two Gymnasium environments. The traffic light chooses between keeping or switching phase (2 actions, with a 15 s minimum green and a 3-step yellow). The vehicle chooses between braking, holding or accelerating (3 actions). Each one observes 6 values.
  • A reward that penalizes waiting, queues and premature switches, and rewards moving vehicles. I tuned it over several iterations.
  • A 3-stage training pipeline with Stable-Baselines3 PPO.
  • An evaluator that runs 6 scenarios, 5 episodes each, against SUMO's fixed-time light.
  • A Three.js dashboard deployed on Netlify.

Architecture

Architecture diagram: traffic counts and the intersection network feed SUMO; a TraCI wrapper connects SUMO to two Gymnasium environments, one for the traffic light and one for the vehicle; each environment trains a PPO agent; an evaluator compares 6 scenarios against the fixed-time light and publishes the results to a Three.js dashboard.

Results

Average waiting time at the intersection, from the project's evaluation:

ScenarioAverage waitImprovement over fixed-time
Fixed-time traffic light303.3 s—
AI traffic light only192.6 s36.5%
AI vehicles only161.9 s46.6%
Full system66.2 s78.2%
Rush hour82.1 s72.9%
Unbalanced traffic78.4 s74.1%

With the full system, throughput dropped from 1,924 to 1,811 vehicles per hour: less waiting, but fewer cars get through.

How to read these numbers. They come from a single deterministic run, since SUMO uses its default seed, and the results CSVs are not in the repository. Before treating them as final, the evaluation needs to be repeated with several seeds after fixing the points in the next section.

Known limits

  • The geometry is synthetic: a 4-arm crossing with 200 m arms and straight-through flows only. The lane counts are real.
  • The environment uses phase index 2 for the Carranza green, but the SUMO network has 8 phases and that green is phase 4. This needs fixing and a new evaluation.
  • The vehicle agent controls 1 vehicle at a time.
  • The two agents were trained separately. The joint fine-tuning stage exists in the pipeline but is not called from main.py.
  • The 35 tests run against a fake simulator and no longer match the environment: they expect a 9-value observation and the environment produces 6. There is no CI.

Links