WAYFINDERRL / 01
Runs on your device
REINFORCEMENT LEARNING OBSERVATORY

Intelligence, in motion.

A world to explore. A policy to discover. Watch an agent learn its own way.

Archipelago

21 × 17 TERRAIN · 210 WALKABLE CELLS
READY TO EXPLORE
3D WORLD · LIVE SIMULATION
DRAG TO ORBIT · RIGHT-DRAG TO PAN · SCROLL TO ZOOM
N
Selected cell: none. Use the coordinate inspector to select or edit a cell.
AgentBeaconHazard -6.0

Current episode · revisits overlap

Orbit, explore, inspect a cell
35%

Select an object, then place it on the terrain. Use Move to reposition existing objects.

Episodes completed 0/ 2,000
Mean reward R

Last 50 episodes

Beacon reached %

Success · last 50 episodes

Exploration 100.0%

0 environment steps

The learning curve

Episode reward20-episode mean
52-2-5EP 0EPISODES →
A learning curve begins with a little exploration.Start training to watch the reward evolve.

Inside the policy

CELL 2, 14
North
0.00
East
0.00
South
0.00
West
0.00
LaunchSelect an object to place it at the coordinates below.
Inspect cell
THE LEARNING LOOP

One decision at a time.

Inspect the latest real transition.

01 / OBSERVEThe worldCurrent state
02 / ACTChoose a moveExplore or exploit
03 / RECEIVEA rewardEnvironment feedback
04 / UPDATERefine a valueTemporal difference error
CURRENT EPISODE 0 stepsEPISODE RETURN 0.00WORLD MODEL 0 state-actionsPLANNING UPDATES 0

What has it really learned?

Evaluate above the map to show the learned path and compare with chance.

WAYFINDER / Made for the joy of figuring things out.
NEXT.JS + THREE.JS