Neural Networks and Deep Learning
Deep learning architectures: artificial neurons, activation functions, backpropagation, CNNs, RNNs, and the dawn of attention mechanisms.
Course Syllabus & Units
Basics
2 lessonsActivation functions
**Activations** are simple nonlinear functions applied between layers so networks can learn curved decision boundaries—not only straight lines.
Neural networks
A **neural network** is a stack of simple math units (neurons) that transform inputs into outputs. With enough layers and data, networks learn useful internal features—edges in images, patterns in sound, structure in text.
Training
2 lessonsBackpropagation
**Backpropagation** is how a neural net figures out **which weights to tweak** after a mistake. Errors flow backward through the layers so each weight gets a “this way / that way” signal.
Optimizers: SGD and Adam
An **optimizer** decides how to nudge weights using gradients. **SGD** steps opposite the gradient (often with minibatches and momentum). **Adam** adapts step sizes using running averages of gradient statistics.
Vision
2 lessonsCNNs and computer vision
**CNNs (convolutional neural networks)** are neural nets designed for grid-like data such as images. They slide small filters across the image to detect local patterns—edges, textures—then build toward objects.
Transfer learning for vision
Reusing a model trained on a large vision dataset as a starting point for your smaller task—**transfer learning** / fine-tuning.
Regularization
1 lessonsModern
3 lessonsAttention outside NLP
How **attention** ideas show up beyond language: vision transformers, multimodal fusion, and set reasoning—while remembering attention ≠ explanation.
Batch norm vs layer norm (deeper)
A deeper comparison of **batch normalization** vs **layer normalization** (and friends): what statistics they use and when each shows up.
Residuals and normalization
**Residual connections** let a layer learn a change to add on top of its input (“skip connections”). **Normalization layers** re-center/re-scale activations so training stays stable.
Practice
5 lessonsData augmentation for vision
Expanding training images with flips, crops, color jitter, and related transforms so models generalize better.
Debugging training curves
Reading **train vs validation loss/accuracy curves** to diagnose underfit, overfit, instability, and data bugs.
Learning rate schedules
Plans for changing the **learning rate** during training: warmups, decays, cosine schedules, plateaus.
Vanishing and exploding gradients
When training deep or recurrent nets, gradients can **shrink toward zero** (vanishing) or **blow up** (exploding) as they pass through many layers/time steps—blocking useful learning.
Weight initialization
How you **set starting weights** before training. Bad init can stall learning; good init keeps signals and gradients in a healthy range.
