Từ GPT-3.5 của ChatGPT (2022) đến Claude 4.7, LLaMA 4, DeepSeek R1 - cách LLM dự đoán token, Transformer, pretrain + SFT + RLHF, Constitutional AI, KV cache, speculative decoding, MoE, reasoning.
23 · iv · 2026
1 posts tagged "rlhf".
Từ GPT-3.5 của ChatGPT (2022) đến Claude 4.7, LLaMA 4, DeepSeek R1 - cách LLM dự đoán token, Transformer, pretrain + SFT + RLHF, Constitutional AI, KV cache, speculative decoding, MoE, reasoning.
23 · iv · 2026