Initializing...
AI development enables advances in robotics. I gave AI agents the task of developing my long-time dream of a tentacle robot.
Three.js · Rapier 3D Physics · Neural Network · Reinforcement Learning (Actor-Critic) · Inverse Kinematics (FABRIK)
Pick ripe (red) strawberries
Avoid unripe (white) ones!
This is a 3D simulation of a tentacle that learns to reach a given target. The tentacle uses a neural network and reinforcement learning to learn smooth movement.
| Click 3D space | Sets the target (red ball) the tentacle tries to reach |
| Drag mouse | Rotates the camera |
| Scroll | Zoom in/out |
| Segments | Number of tentacle joints/segments. More = longer and more flexible, but harder to learn. |
| Learning Rate | Learning speed. Higher = faster learning but less stable. Lower = more stable but slower. |
| Reset Tentacle | Resets the tentacle to straight position (does not reset learning) |
| Reset Brain | Resets the neural network and starts learning from scratch (with pretraining) |
| Enable Obstacles | Activates obstacles. The tentacle must learn to avoid them to reach the target. |
| Obstacles | Number of obstacles (1-10). More obstacles = harder task. |
| Regenerate | Creates new random obstacles (spheres, cylinders, boxes). |
Tip: Start without obstacles, let the tentacle learn basic movement. Then gradually add obstacles.
| Episode | One attempt to reach the target. Ends when the target is reached or time runs out. Success rate shown in parentheses. |
| Distance | Distance from tentacle tip to target. Lower = better. |
| Avg Reward | Average reward from recent episodes. Increasing value = learning is progressing. |
| Neural Network (NN) | Artificial "brain" that learns to map situations (where the target is) to actions (how the joints move). |
| Reinforcement Learning | A learning method where the agent (tentacle) learns by trial and error. Gets rewards for good actions, penalties for bad ones. |
| Actor-Critic | Two neural networks: the Actor decides actions, the Critic evaluates how good a situation is. |
| FABRIK (IK) | Inverse Kinematics algorithm that calculates joint angles to reach a target. Used as prior knowledge. |
| Exploration | Random experimentation. High at first (learns new things), decreases over time (exploits what was learned). |
| Policy Gradient | Learning algorithm that reinforces actions that led to good outcomes. |
| Pretraining | Initial training with the IK algorithm. Gives the tentacle an "instinct" to reach toward targets before actual learning begins. |
A small view in the bottom-right corner shows what the tentacle tip sees. This simulates "vision" that the tentacle can use to find targets.
| Tip Camera | Shows/hides the tip camera view. |
| Strategy | Selects how the tentacle uses the camera: |
Indicators: Green = target visible, Red = searching. Crosshair shows the camera center point.
Tentacle Simulation v1.0
Neural network + Reinforcement Learning + Three.js