Resources
Papers, documentation and standards
The original sources and reference manuals behind the course, gathered in one place.
Research papers
The original papers behind ideas taught in the course.
- Attention Is All You Need (Vaswani et al., 2017): The paper that introduced the Transformer.
- LoRA: Low-Rank Adaptation of Large Language Models (Hu et al., 2021): Cheap fine-tuning with small low-rank matrices.
- Prefix-Tuning (Li and Liang, 2021): Learning a short task-specific prefix while the model stays frozen.
- Distilling the Knowledge in a Neural Network (Hinton, Vinyals and Dean, 2015): Knowledge distillation from a teacher to a student model.
- Deep Reinforcement Learning from Human Preferences (Christiano et al., 2017): Learning a reward model from human comparisons.
- Training Language Models to Follow Instructions with Human Feedback (Ouyang et al., 2022): The InstructGPT paper: RLHF applied to language models.
- Proximal Policy Optimization Algorithms (Schulman et al., 2017): PPO, the policy-gradient method used in RLHF.
- Direct Preference Optimization (Rafailov et al., 2023): Preference alignment without a separate reward model.
- DeepSeekMath (introduces GRPO) (Shao et al., 2024): Group Relative Policy Optimization.
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (Lewis et al., 2020): The paper that named RAG.
- RoFormer: Rotary Position Embedding (Su et al., 2021): RoPE, the positional scheme used by many modern LLMs.
Official documentation
Reference manuals for the tools used in the code examples.
- Python documentation: The language used in every code example and in the Practice editor.
- NumPy documentation: Arrays and linear algebra.
- PyTorch documentation: The deep-learning framework most examples refer to.
- scikit-learn user guide: Classical machine-learning algorithms and metrics.
- Hugging Face Transformers: Loading, running and fine-tuning Transformer models.
- vLLM documentation: High-throughput LLM serving.
Standards and specifications
Open formats covered in the agents modules.
- Model Context Protocol: The open protocol for connecting models to tools and data.
- Agent Skills: The open specification for packaging agent skills.