Glossary

Reinforcement learning (RL)

Beginner

Learning by trial and error: a program tries actions, gets a score, and slowly learns which choices lead to better scores.

Novice

A program (the agent) looks at a situation, takes an action and gets a score (the reward). Over many tries it learns a strategy that earns a high total score, for example where to put each large block on a chip.

Expert

Formally a Markov decision process: states, actions, the rules for moving between states, and a reward. In chip design the true score needs hours of placement and routing, so agents train on a fast estimate instead, and the result is only as good as that estimate’s agreement with the finished layout.

Explained in AI in EDA (Extras).

See also: Proxy metric.

All 896 terms →