IT lexicon AI & ML OpenAI o1

OpenAI o1

AI & ML På svenska → Updated: 2026-05-23

OpenAI's first reasoning model (September 2024) — emits an internal "thinking" CoT phase before answering and proved that test-time compute can beat raw parameter scaling.

Trained via RL where the reward signal is correct final answers on maths and code problems. The model learns to produce long internal reasoning traces (partly hidden from the user). Set new SOTA on AIME, USAMO and Codeforces by a wide margin.

Triggered the entire "reasoning model" wave: DeepSeek-R1, QwQ, Gemini Thinking, Claude Extended Thinking, GPT o3/o4. Trade-off: expensive and slow — answers take seconds to minutes.

← Back to the lexicon