Reinforcement learning · live in your browser

Teaching a traffic light to think ahead.

A single intersection, two queues, one decision every tick: which road gets the green? Tabular Q-learning and Monte Carlo agents learn it from experience, and are checked against classic signal timing and the mathematically optimal policy.

Q-learning vs fixed-time signals–fewer cars waiting
Gap to the optimal policy–avg cars waiting per step
States matching optimal–Q-learning greedy policy
Building the intersection…
Q-learning
tick 0
East–West green
North–South0/3
East–West0/5
Drag to orbit · scroll to zoom

Ask the agent

Set the queues and call the live /api/decide endpoint. The state is also loaded into the 3D scene.

1
3

Benchmark in your browser

Runs every controller on the same 400 random 50-tick episodes (common random numbers) and ranks them.

Results

How well did the agents learn?

5,000 training episodes, 5 independent seeds per algorithm, evaluated on 2,000 held-out traffic episodes that every controller sees identically. Lower is better.

ControllerAvg cars waiting95% CITurned away / tickMatches optimal
Bar chart of average cars waiting per controller
Evaluation against baselines and the exact optimum.
Learning curves of Q-learning and Monte Carlo
Learning curves (mean of 5 seeds, min–max band).
Heatmaps of greedy actions per state
What each agent does in every one of the 24 states.

Finding: the reward lets agents starve the short road

Cars that arrive at a full queue simply leave the model, and the original reward never counts them. So the best-scoring strategy is to keep East–West green almost always: the 3-car North–South queue sits full and extra cars are turned away. The overflow-aware variant charges 2 points per turned-away car; its optimal policy serves North–South more often and turns away far fewer cars, at the cost of slightly longer queues. You can watch both in the simulator.

Hyperparameter experiments

RunAlgorithmαγε₀ → minε decayEpisodesAvg waiting (± sd)Matches optimal
How it works

The problem, as a Markov decision process

◎

State

Queue lengths (ns, ew): 0–3 cars North–South, 0–5 East–West. 24 states.

⇄

Action

Give the green to East–West or North–South for the next tick.

⤳

Dynamics

0–2 cars arrive on each road; the green road discharges 1–2. Queues are capped.

−Σ

Reward

−(ns + ew) each tick: every waiting car costs a point. Episodes last 50 ticks.

Q-learning

Off-policy temporal-difference control. After every tick it nudges Q(s,a) toward r + γ·max Q(s′,·), with ε-greedy exploration decaying from 1.0 to 0.01.

Monte Carlo control

On-policy, first-visit. Plays a whole episode, then averages the actual discounted return that followed each first visit of a state–action pair.

Value iteration

The state space is small enough to solve exactly: the full transition model gives the optimal policy, the ground truth both learners are judged against.

API

Use the agents from code

# greedy decision for a state
curl "/api/decide?ns=2&ew=4&policy=q_learning"

# JSON body works too
curl -X POST /api/decide \
  -H "Content-Type: application/json" \
  -d '{"ns": 3, "ew": 1, "policy": "monte_carlo"}'
GET /api/decidens 0–3, ew 0–5, policy: q_learning · monte_carlo · optimal · optimal_overflow_aware
GET /api/policiesAll Q-tables as JSON
GET /api/resultsEvaluation results
GET /api/healthLiveness check

Invalid input returns HTTP 400 with a readable error message.

Team

Built at Dublin Business School

Reinforcement Learning module · CA2 group project: “Optimising traffic light decisions to reduce waiting time and congestion”. Read the original report →

AIAloke Kunjandi Issac
ASAnamika Koonakkampilly Sunilal
DJDeninson David Jeyakumar
MFMuhammed Fayiz Kizhakkumparamban
IJIrene Ann Jacob