Teaching a traffic light to think ahead.
A single intersection, two queues, one decision every tick: which road gets the green? Tabular Q-learning and Monte Carlo agents learn it from experience, and are checked against classic signal timing and the mathematically optimal policy.
Ask the agent
Set the queues and call the live /api/decide endpoint. The state is also loaded into the 3D scene.
Benchmark in your browser
Runs every controller on the same 400 random 50-tick episodes (common random numbers) and ranks them.
How well did the agents learn?
5,000 training episodes, 5 independent seeds per algorithm, evaluated on 2,000 held-out traffic episodes that every controller sees identically. Lower is better.
| Controller | Avg cars waiting | 95% CI | Turned away / tick | Matches optimal |
|---|



Finding: the reward lets agents starve the short road
Cars that arrive at a full queue simply leave the model, and the original reward never counts them. So the best-scoring strategy is to keep East–West green almost always: the 3-car North–South queue sits full and extra cars are turned away. The overflow-aware variant charges 2 points per turned-away car; its optimal policy serves North–South more often and turns away far fewer cars, at the cost of slightly longer queues. You can watch both in the simulator.
Hyperparameter experiments
| Run | Algorithm | α | γ | ε₀ → min | ε decay | Episodes | Avg waiting (± sd) | Matches optimal |
|---|
The problem, as a Markov decision process
State
Queue lengths (ns, ew): 0–3 cars North–South, 0–5 East–West. 24 states.
Action
Give the green to East–West or North–South for the next tick.
Dynamics
0–2 cars arrive on each road; the green road discharges 1–2. Queues are capped.
Reward
−(ns + ew) each tick: every waiting car costs a point. Episodes last 50 ticks.
Q-learning
Off-policy temporal-difference control. After every tick it nudges Q(s,a) toward r + γ·max Q(s′,·), with ε-greedy exploration decaying from 1.0 to 0.01.
Monte Carlo control
On-policy, first-visit. Plays a whole episode, then averages the actual discounted return that followed each first visit of a state–action pair.
Value iteration
The state space is small enough to solve exactly: the full transition model gives the optimal policy, the ground truth both learners are judged against.
Use the agents from code
# greedy decision for a state curl "/api/decide?ns=2&ew=4&policy=q_learning" # JSON body works too curl -X POST /api/decide \ -H "Content-Type: application/json" \ -d '{"ns": 3, "ew": 1, "policy": "monte_carlo"}'
GET /api/decide | ns 0–3, ew 0–5, policy: q_learning · monte_carlo · optimal · optimal_overflow_aware |
GET /api/policies | All Q-tables as JSON |
GET /api/results | Evaluation results |
GET /api/health | Liveness check |
Invalid input returns HTTP 400 with a readable error message.
Built at Dublin Business School
Reinforcement Learning module · CA2 group project: “Optimising traffic light decisions to reduce waiting time and congestion”. Read the original report →