IT lexicon AI & ML Reinforcement learning

Reinforcement learning RL

AI & ML På svenska → Updated: 2026-05-23

ML paradigm where an agent learns through reward — try, get feedback, improve.

Behind AlphaGo, robotics, self-driving. Drove RLHF as the finish on every modern LLM. Tricky to design the right reward function — wrong reward, wrong behaviour.

← Back to the lexicon