Course Overview
This course explains the fundamental principles behind large language models (LLMs) like ChatGPT by tracing the evolution of language modeling techniques from 1913 to 2026. It covers key models and architectures such as N-grams, RNNs, LSTMs, Attention mechanisms, Transformers, Mamba, and Gated DeltaNet, focusing on the problems each innovation solved. The course uses minimal math and programming, relying on analogies and clear explanations to make complex concepts accessible.
Key Takeaways
- Understand that language models predict the probability of the next word and that text generation is a repeated next-word guessing game.
- Learn how N-grams, RNNs, LSTMs, Attention, Transformers, Mamba, and Gated DeltaNet each address specific challenges in language modeling.
- Manually compute N-gram probabilities and Self-Attention with simple numbers to grasp their mechanics.
- Comprehend terms like embedding, gate, KV cache, state space model, MoE, and attention sink through analogies.
- Compare Transformer and Mamba architectures, understanding their strengths and weaknesses and the rationale for hybrid models.
- Interpret the hybrid architecture of Qwen3-Next / 3.5 / 3.8 combining multiple Gated DeltaNet and Gated Attention layers.
Prerequisites
- No prior math or programming knowledge required; basic addition, multiplication, and proportional reasoning suffice.
- Experience using conversational AI tools like ChatGPT is helpful but not mandatory.
- Access to a browser with Google Colab is useful for optional hands-on exercises.
Target Learners
- Users of AI tools who want to understand the inner workings of language models.
- Individuals intimidated by technical terms in LLM articles and news.
- Beginners seeking a systematic, non-technical introduction to AI principles.
- Those interested in reading and contextualizing specifications and research papers on modern language models.
- Full Pack
Discussions are closed
Comments are currently disabled for this course.
