Hide & Seek RL Lab
Runs only on this device

Booting

Training

Live simulationPreparation
0% visible

Rounds in this run

4 policy networks · 357,258 parameters each · data stays in your browser

200rounds
100350
Estimated timeInitializing neural networks

SELF-PLAY POPULATION

Policy population

Not saved yet
Network 01Hider policy
ELO1000
0 wins357.3k params
Network 02Seeker policy
ELO1000
0 wins357.3k params
Network 03Hider policy
ELO1000
0 wins357.3k params
Network 04Seeker policy
ELO1000
0 wins357.3k params
Policy loss0.000
Value loss0.000
Policy entropy2.480
CheckpointIndexedDB