AI 4200

Deep Learning

Today

  1. 01Syllabus has changed
  2. 02Quiz review
  3. 03Multi-class classification
  4. 04Optimizers
  5. 05Learning rate scheduling
  6. 06Dropout
  7. 07Project 1

Change to the syllabus

  • Based on discussion with the class and a vote
  • Exams
    • Previously 50% of grade
      • 25% for midterm
      • 25% for final
    • Now Final exam is 20%
      • Will be based on the quizzes
  • Projects
    • Previously 0%
    • Now 3-ish projects totaling 30% of grade
  • Quizzes and homeworks are unchanged

Quiz review

  1. Backprop
    1. Chain rule
    2. Transpose before final multiplication with inputs to layer to get shape of weights
  2. Neural nets with linear activation collapse to a single layer linear model
  3. A vanishes, B is stable and usable, C explodes, D is stable but zero
  4. \(E[L_2^2] = N\sigma^2\), important because if means the variance of the gradient directly relates to the magnitude of the gradient
  5. Variance grows (exploding grad), variance shrinks (vanishing grad), stable
  6. Output of sigmoid is option c (mean 0.5 range 0 to 1), tanh is zero mean
  7. \(\hat{Y} = XW^T\), \(X\) is eXf, \(W\) is nXf, \(\hat{Y}\) is eXn
  8. \(\sigma_{\hat{y}}^2 = f\sigma_w^2\sigma_x^2\)
  9. \(\sigma_w^2 = \frac{1}{f},\;\; \sigma_{\hat{y}}^2 = \sigma_x^2\) (Bengio)
  10. \(\frac{\partial L}{\partial \hat{Y}}\) is eXn, \(\;Var\!\left(\frac{\partial L}{\partial \hat{Y}}\right) = \sigma_g^2,\;\; \frac{\partial L}{\partial X} = \frac{\partial L}{\partial \hat{Y}}\frac{\partial \hat{Y}}{\partial X} = \frac{\partial L}{\partial \hat{Y}}W,\;\; Var\!\left(\frac{\partial L}{\partial X}\right) = n\sigma_g^2\sigma_w^2,\;\; \sigma_w^2 = \frac{1}{n}\)
    1. Glorot compromise \(\sigma_w^2 = \frac{2}{f+n}\)
  11. \(E[H^2] = ReLU\big(E[\hat{Y}^2]\big) = \frac{1}{2}E[\hat{Y}^2] = \frac{\sigma_{\hat{y}}^2}{2} = \frac{f\sigma_w^2\sigma_x^2}{2},\;\; \sigma_w^2 = \frac{2}{f},\;\; \sigma_{\hat{y}}^2 = \sigma_x^2\) (Kaiming)
  12. Mean of half normal is \(E[H] = \mu = \frac{1}{\sqrt{2\pi}} \approx 0.4\)
  13. \(E[H^2] = 0.5 = \sigma^2 + \mu^2,\;\; 0.5 = 0.4^2 + \sigma^2,\;\; \sigma^2 = 0.5 - 0.4^2 = 0.34\)
  14. ReLU’s second moment is 32% made of the mean, each ReLU layer decreases the amount of input dependent variance in what is passed forward by a factor of 0.68
    1. First moment \(E[H] = \mu\)
    2. Second moment \(E[H^2] = \sigma^2 + \mu^2\)

Multi-class classification

One handwritten digit in, ten classes out. The network produces one raw score — a logit — per digit.

MNIST · 28 × 28 = 784 pixels
Network
784 → … → 10
ŷ = f(x)
LOGITS ŷ · ONE PER DIGIT
0
−1.2
1
0.3
2
1.1
3
0.8
4
−0.5
5
−0.9
6
−2.0
7
4.1
8
0.2
9
1.6
Logits are unbounded and don’t sum to anything — the highest one (7) wins, but they are not yet probabilities.

Why Multi-Class Classification?

  • Couldn’t we just have 10 individual sigmoid outputs
    • Train each with BCE
    • Yes… sometimes it may say
      • 0.9 a 1
      • 0.8 a 7
    • Who cares?
    • That can work – called multi-label network
  • Why do we want output probs that sum to 1?
    • Each image is exactly 1 digit
      • Not multiple digits simultaneously
  • Models should match the inductive bias of the task
    • We will return to this point soon
    • Eliminates bad solutions like the one above

Multi-Class classification with an MLP on MNIST

  • Input size is 28*28 (one for every pixel)
  • Output size is the number of classes (10 because 10 digits)
  • The difference lies only in the output layer
  • A new activation function called softmax
    • An activation layer
    • Takes in all the linear outputs at the last layer
    • Computes activation over all and forces sum to 1
  • ŷ₀ = −1.2
    ŷ₇ = 4.1
    ŷ₉ = 1.6
    e
    e
    e
    eŷ₀ = 0.30
    eŷ₇ = 60.34
    eŷ₉ = 4.95
    Σk eŷk = 74.54
    ÷
    ÷
    ÷
    p̂₀ = 0.004
    p̂₇ = 0.810
    p̂₉ = 0.066
    Softmax forward pass
    Exponentiate every logit, add them up once, then divide each exponential by that one shared sum.
    LOGITS
    SOFTMAX LAYER
    PROBABILITIES
    ⋮
    ⋮
    ⋮
    ⋮
    ⋮
    ⋮
    ⋮
    ⋮
    one sum, shared by all 10 divisions
    All ten p̂ are positive and sum to 1.

    Derivative vs Gradient vs Jacobian

    • Partial Derivative
      • How a scalar output function changes with a scalar input change
      • \(\frac{\partial \widehat{p_2}}{\partial y_1} = \widehat{p_2}\widehat{p_1}\)
    • Gradient
      • How a scalar output function changes with a vector input change
        • \(\frac{\partial p_2}{\partial y} = \left\langle \frac{\partial \widehat{p_2}}{\partial y_1}, \frac{\partial \widehat{p_2}}{\partial y_2}, \frac{\partial \widehat{p_2}}{\partial y_3} \right\rangle = \left\langle \widehat{p_2}\widehat{p_1},\; \widehat{p_2}(1-\widehat{p_2}),\; \widehat{p_2}\widehat{p_3} \right\rangle\)
    • Jacobian
      • How a vector output function changes with a vector input change
      • \(J_p = \frac{\partial \hat{p}}{\partial y} = \begin{bmatrix} \frac{\partial \widehat{p_1}}{\partial y_1} & \frac{\partial \widehat{p_1}}{\partial y_2} & \frac{\partial \widehat{p_1}}{\partial y_3} \\[4pt] \frac{\partial \widehat{p_2}}{\partial y_1} & \frac{\partial \widehat{p_2}}{\partial y_2} & \frac{\partial \widehat{p_2}}{\partial y_3} \\[4pt] \frac{\partial \widehat{p_3}}{\partial y_1} & \frac{\partial \widehat{p_3}}{\partial y_2} & \frac{\partial \widehat{p_3}}{\partial y_3} \end{bmatrix} = \begin{bmatrix} \widehat{p_1}(1-\widehat{p_1}) & \widehat{p_1}\widehat{p_2} & \widehat{p_1}\widehat{p_3} \\ \widehat{p_2}\widehat{p_1} & \widehat{p_2}(1-\widehat{p_2}) & \widehat{p_2}\widehat{p_3} \\ \widehat{p_3}\widehat{p_1} & \widehat{p_3}\widehat{p_2} & \widehat{p_3}(1-\widehat{p_3}) \end{bmatrix}\)
    ∂L/∂ŷ₀
    ∂L/∂ŷ₁
    ∂L/∂ŷ₉
    Softmax Jacobian
    ∂P̂/∂Ŷ
    n × n per example
    mixes every row
    ∂L/∂p̂₀
    ∂L/∂p̂₁
    ∂L/∂p̂₉
    Backprop through softmax needs the Jacobian
    To pass the gradient from the probabilities back to the logits, it has to go through ∂P̂/∂Ŷ.
    LOGIT GRADIENT
    LOSS GRADIENT
    ⋮
    ⋮
    to the linear layer
    from the loss
    backward pass ←
    ∂L/∂Ŷ = ∂L/∂P̂ · ∂P̂/∂Ŷ
    ∂L/∂P̂
    Comes from the loss — softmax doesn’t care which one. Shape: examples × classes.
    ∂P̂/∂Ŷ
    A full classes × classes matrix for every example: each p̂ᵢ depends on every ŷⱼ.
    ∂L/∂Ŷ = ∂L/∂P̂ · ∂P̂/∂Ŷ
    A vector–Jacobian product per example. Shape: examples × classes — ready for the linear layer’s ∂L/∂W.
    Compare sigmoid: ∂P̂/∂Ŷ = P̂(1 − P̂) is diagonal, so the product collapses to an element-wise (Hadamard) multiply. Softmax’s ∂P̂/∂Ŷ isn’t.

    Building the Jacobian costs O(n²)

    Every logit sits in the shared denominator, so every output depends on every input.

    Sigmoid∂P̂/∂Ŷn entries
    ŷ₀ŷ₁ŷ₂ŷ₃ŷ₄ŷ₅ŷ₆ŷ₇ŷ₈ŷ₉
    p̂₀p̂₁p̂₂p̂₃p̂₄p̂₅p̂₆p̂₇p̂₈p̂₉
    p̂₀ (1−p̂₀)
    0
    0
    0
    0
    0
    0
    0
    0
    0
    0
    p̂₁ (1−p̂₁)
    0
    0
    0
    0
    0
    0
    0
    0
    0
    0
    p̂₂ (1−p̂₂)
    0
    0
    0
    0
    0
    0
    0
    0
    0
    0
    p̂₃ (1−p̂₃)
    0
    0
    0
    0
    0
    0
    0
    0
    0
    0
    p̂₄ (1−p̂₄)
    0
    0
    0
    0
    0
    0
    0
    0
    0
    0
    p̂₅ (1−p̂₅)
    0
    0
    0
    0
    0
    0
    0
    0
    0
    0
    p̂₆ (1−p̂₆)
    0
    0
    0
    0
    0
    0
    0
    0
    0
    0
    p̂₇ (1−p̂₇)
    0
    0
    0
    0
    0
    0
    0
    0
    0
    0
    p̂₈ (1−p̂₈)
    0
    0
    0
    0
    0
    0
    0
    0
    0
    0
    p̂₉ (1−p̂₉)
    Softmax∂P̂/∂Ŷn² entries
    ŷ₀ŷ₁ŷ₂ŷ₃ŷ₄ŷ₅ŷ₆ŷ₇ŷ₈ŷ₉
    p̂₀p̂₁p̂₂p̂₃p̂₄p̂₅p̂₆p̂₇p̂₈p̂₉
    p̂₀ (1−p̂₀)
    −p̂₀p̂₁
    −p̂₀p̂₂
    −p̂₀p̂₃
    −p̂₀p̂₄
    −p̂₀p̂₅
    −p̂₀p̂₆
    −p̂₀p̂₇
    −p̂₀p̂₈
    −p̂₀p̂₉
    −p̂₁p̂₀
    p̂₁ (1−p̂₁)
    −p̂₁p̂₂
    −p̂₁p̂₃
    −p̂₁p̂₄
    −p̂₁p̂₅
    −p̂₁p̂₆
    −p̂₁p̂₇
    −p̂₁p̂₈
    −p̂₁p̂₉
    −p̂₂p̂₀
    −p̂₂p̂₁
    p̂₂ (1−p̂₂)
    −p̂₂p̂₃
    −p̂₂p̂₄
    −p̂₂p̂₅
    −p̂₂p̂₆
    −p̂₂p̂₇
    −p̂₂p̂₈
    −p̂₂p̂₉
    −p̂₃p̂₀
    −p̂₃p̂₁
    −p̂₃p̂₂
    p̂₃ (1−p̂₃)
    −p̂₃p̂₄
    −p̂₃p̂₅
    −p̂₃p̂₆
    −p̂₃p̂₇
    −p̂₃p̂₈
    −p̂₃p̂₉
    −p̂₄p̂₀
    −p̂₄p̂₁
    −p̂₄p̂₂
    −p̂₄p̂₃
    p̂₄ (1−p̂₄)
    −p̂₄p̂₅
    −p̂₄p̂₆
    −p̂₄p̂₇
    −p̂₄p̂₈
    −p̂₄p̂₉
    −p̂₅p̂₀
    −p̂₅p̂₁
    −p̂₅p̂₂
    −p̂₅p̂₃
    −p̂₅p̂₄
    p̂₅ (1−p̂₅)
    −p̂₅p̂₆
    −p̂₅p̂₇
    −p̂₅p̂₈
    −p̂₅p̂₉
    −p̂₆p̂₀
    −p̂₆p̂₁
    −p̂₆p̂₂
    −p̂₆p̂₃
    −p̂₆p̂₄
    −p̂₆p̂₅
    p̂₆ (1−p̂₆)
    −p̂₆p̂₇
    −p̂₆p̂₈
    −p̂₆p̂₉
    −p̂₇p̂₀
    −p̂₇p̂₁
    −p̂₇p̂₂
    −p̂₇p̂₃
    −p̂₇p̂₄
    −p̂₇p̂₅
    −p̂₇p̂₆
    p̂₇ (1−p̂₇)
    −p̂₇p̂₈
    −p̂₇p̂₉
    −p̂₈p̂₀
    −p̂₈p̂₁
    −p̂₈p̂₂
    −p̂₈p̂₃
    −p̂₈p̂₄
    −p̂₈p̂₅
    −p̂₈p̂₆
    −p̂₈p̂₇
    p̂₈ (1−p̂₈)
    −p̂₈p̂₉
    −p̂₉p̂₀
    −p̂₉p̂₁
    −p̂₉p̂₂
    −p̂₉p̂₃
    −p̂₉p̂₄
    −p̂₉p̂₅
    −p̂₉p̂₆
    −p̂₉p̂₇
    −p̂₉p̂₈
    p̂₉ (1−p̂₉)
    Row i: output p̂ᵢ
    Column j: logit ŷⱼ
    Cell: ∂p̂ᵢ/∂ŷⱼ
    p̂i = eŷi / Σk eŷk
    Nudge any ŷj and the sum changes, so every p̂i moves:
    ∂p̂i/∂ŷj = p̂i(δij − p̂j)
    ENTRIES PER EXAMPLE
    MNIST10²ImageNet1,000² = 10⁶50k vocab50,000² = 2.5 × 10⁹
    positive negative exactly 0
    In general, off-diagonal Jacobian entries are the problem: they couple the outputs, and coupling is what makes the matrix n × n.

    The product is cheaper than the matrix

    Backprop never needs ∂P̂/∂Ŷ itself — only the product ∂L/∂P̂ · ∂P̂/∂Ŷ. And every entry of that product reuses one sum.

    ∂p̂ᵢ/∂ŷⱼ = p̂ᵢ(δᵢⱼ − p̂ⱼ) = p̂ᵢδᵢⱼ − p̂ᵢp̂ⱼ δᵢⱼ is 1 only when i = j: the blue term sits on the diagonal, the orange term fills every cell.
    p̂₀(1−p̂₀)
    −p̂₀p̂₁
    −p̂₀p̂₂
    −p̂₁p̂₀
    p̂₁(1−p̂₁)
    −p̂₁p̂₂
    −p̂₂p̂₀
    −p̂₂p̂₁
    p̂₂(1−p̂₂)
    ∂P̂/∂Ŷ
    =
    p̂₀
    0
    0
    0
    p̂₁
    0
    0
    0
    p̂₂
    diag(p̂)
    −
    p̂₀p̂₀
    p̂₀p̂₁
    p̂₀p̂₂
    p̂₁p̂₀
    p̂₁p̂₁
    p̂₁p̂₂
    p̂₂p̂₀
    p̂₂p̂₁
    p̂₂p̂₂
    p̂ p̂ᵀ (outer product)
    Check a diagonal cell: p̂₀ − p̂₀p̂₀ = p̂₀(1−p̂₀). Off the diagonal only −p̂ᵢp̂ⱼ is left.
    EACH ENTRY OF ∂L/∂P̂ · ∂P̂/∂Ŷ, WRITTEN OUT (n = 3)
    ∂L/∂ŷ₀ =
    p̂₀ ∂L/∂p̂₀
    − p̂₀ ·( p̂₀ ∂L/∂p̂₀ + p̂₁ ∂L/∂p̂₁ + p̂₂ ∂L/∂p̂₂ )
    Orange: from p̂p̂ᵀ.
    Dashed: identical in every row — compute it once as s
    ∂L/∂ŷ₁ =
    p̂₁ ∂L/∂p̂₁
    − p̂₁ ·( p̂₀ ∂L/∂p̂₀ + p̂₁ ∂L/∂p̂₁ + p̂₂ ∂L/∂p̂₂ )
    ∂L/∂ŷ₂ =
    p̂₂ ∂L/∂p̂₂
    − p̂₂ ·( p̂₀ ∂L/∂p̂₀ + p̂₁ ∂L/∂p̂₁ + p̂₂ ∂L/∂p̂₂ )
    s = Σᵢ p̂ᵢ ∂L/∂p̂ᵢn multiplies, done once
    ∂L/∂ŷⱼ = p̂ⱼ(∂L/∂p̂ⱼ − s)O(1) per entry
    O(n)
    Before building a Jacobian with off-diagonal entries, look for shared structure — here a common sum turns n² work into two linear passes.

    From BCE to cross-entropy

    The only difference is how many jobs each output has.

    Binary · one outputsigmoid p̂
    low p̂ → class 0
    class 1 ← high p̂
    00.51
    Both ends of the same number are an answer. The output does double duty, so the likelihood needs a term for each meaning:
    Pr(label) = p̂p · (1−p̂)1−p
    BCE = −[ p log p̂ + (1−p) log(1−p̂) ]
    Multi-class · one output per classsoftmax p̂
    0123456789
    A low p̂₃ only means “not a 3” — it doesn’t point to any other class. Each output has one job, so the second term disappears:
    Pr(label) = Πk p̂kpk = p̂true
    CE = −Σk pk log p̂k
    Change the (1−p̂)1−p term to a separate output and what’s left is −p log p̂ summed over classes.
    cross-entropy is BCE.

    Cross-entropy loss

    The negative log-likelihood that the model assigns every input the correct label in the dataset.

    label 7
    0.81
    ×
    label 2
    0.93
    ×
    label 1
    0.97
    ×
    label 0
    0.64
    = 0.47
    Probability every label is correct (examples independent)
    Pr = Πi p̂i,true = 0.81 × 0.93 × 0.97 × 0.64 = 0.47
    Take −log: the product becomes a sum
    −log Pr = −Σi log p̂i,true
           = 0.21 + 0.07 + 0.03 + 0.45 = 0.76
    Average over the N examples → L = 0.76 / 4 = 0.19
    Per example, against the actual distribution p
    H(p, p̂) = −Σk pk log p̂k

    p is the true probability of each class. For MNIST it’s one-hot (1 on the correct digit, 0 elsewhere)

    so only one term survives: −log p̂true.

    Minimizing cross-entropy = maximizing the probability the model assigns to the true labels.

    CrossEntropyLoss takes logits

    1PyTorch folds softmax into the loss: nn.CrossEntropyLoss = log-softmax + negative log-likelihood.
    2Numerically stable: log p̂ⱼ = ŷⱼ − logsumexp(ŷ), computed with the max subtracted. No overflow, no log(0), no clipping.
    3The model now returns logits, so inference needs a predict method that applies softmax to get probabilities back.
    And ∂P̂/∂Ŷ disappears. With ∂L/∂p̂ⱼ = −pⱼ/p̂ⱼ:
    s = Σᵢ p̂ᵢ ∂L/∂p̂ᵢ = −Σ pᵢ = −1
    p̂ⱼ(∂L/∂p̂ⱼ − s) = p̂ⱼ(1 − pⱼ/p̂ⱼ)
    ∂L/∂Ŷ = P̂ − P
    Softmax in the model
    class MLP(nn.Module):
        # __init__: fc1 = Linear(784, 128), fc2 = Linear(128, 10)
    
        def forward(self, x):
            h = torch.tanh(self.fc1(x))
            p_hat = torch.softmax(self.fc2(h), dim=1)
            return p_hat.clamp(1e-7, 1.0)    # avoid log(0)
    
    p_hat = model(x)
    loss = nn.NLLLoss()(torch.log(p_hat), labels)
    PyTorch way
    class MLP(nn.Module):
        # __init__: fc1 = Linear(784, 128), fc2 = Linear(128, 10)
    
        def forward(self, x):
            h = torch.tanh(self.fc1(x))
            return self.fc2(h)             # raw logits
    
        def predict(self, x):
            return torch.softmax(self.forward(x), dim=1)  # p̂
    
    loss = nn.CrossEntropyLoss()(model(x), labels)
    Softmax now lives inside the loss, so it’s gone at inference — predict puts it back, returning the full distribution p̂ for confidence, top-k or argmax.

    Optimizers change progress through the loss landscape

    • SGD
      • The base approach
      • Flag for momentum (exponential moving average of gradient) \(m\)
        • \(v_{t+1} = m v_t + \frac{\partial L}{\partial W},\quad W_{t+1} = W_t - \alpha v_{t+1}\)
      • Flag for weight decay \(\lambda\)
        • \(W_{t+1} = (1-\alpha\lambda)W_t - \alpha v_{t+1}\)
    • Adaptive extensions
      • Adagrad
        • SGD scaled by the L2 norm of ALL past updates for each weight
      • RMSProp
        • SGD scaled by RMS (second moment) over a recent window of each weight’s gradient
      • AdamW
        • Combines RMSProp with momentum
        • Forces each weight to update relatively equally considering gradient landscape differences
    • Newton methods
      • LBFGS – memory hungry due to second derivative jacobian non-diagonal
    3D surface plot of a loss landscape titled Gradient-decay collapse. Two deep valleys mark good minima, separated by a steep ridge; between them a shallow cove marks the trivial mean solution. A ball sits in the cove because it cannot climb the walls to reach a valley, so it settles in the worse minimum.

    Pytorch

    import torch
    import torch.nn as nn
    import torch.optim as optim
    import torch.optim.lr_scheduler as lr_scheduler
    # 1. Define your model, loss, and optimizer
    model = nn.Linear(10, 2)
    criterion = nn.MSELoss()
    optimizer = optim.SGD(model.parameters(), lr=0.01, momentum=0.9)
    optimizer = optim.Adagrad(model.parameters(), lr=0.01)
    optimizer = optim.RMSprop(model.parameters(), lr=0.01, alpha=0.99, momentum=0.9)
    optimizer = torch.optim.AdamW( model.parameters(), lr=1e-3, weight_decay=0.01 )
    optimizer = optim.LBFGS(model.parameters(), lr=1.0)

    Fixed learning rates oscillate

    Two parabolas comparing fixed learning rates. With learning rate 0.5 the iterates bounce in a wide cloud around the minimum: typical distance about 0.58, average loss floor about 0.17. With learning rate 0.05 the cloud is tight: typical distance about 0.16, floor about 0.013. Every step is a tug of war between the pull toward the minimum (learning rate times w) and a random kick (learning rate times noise), giving a floor of roughly learning rate times noise size squared over 4.

    Fixed learning rates are fast or precise

    Log-scale loss over 500 steps for two fixed learning rates. Learning rate 0.5 drops to its floor in about 10 steps and then stalls at 0.17. Learning rate 0.01 keeps improving slowly and is still approaching a floor about 65 times lower at step 500. Early in training a large learning rate gives speed; late in training a small one gives precision.

    Variable learning rates that are scheduled

    Two stacked plots. Top: learning rate over 500 steps for a step-decay schedule (divided by 10 at steps 60 and 200) and a cosine schedule that decays smoothly to 0. Bottom: log-scale loss. Fixed learning rate 0.5 stalls high and fixed 0.01 improves slowly, while both schedules fall as fast as the large rate at first and then keep dropping, since each tenfold cut in learning rate lowers the floor tenfold. A schedule gets the speed of a large learning rate and the precision of a small one.
    Learning rate
    Loss

    Pytorch

    import torch
    import torch.nn as nn
    import torch.optim as optim
    import torch.optim.lr_scheduler as lr_scheduler
    # 1. Define your model, loss, and optimizer
    model = nn.Linear(10, 2)
    criterion = nn.MSELoss()
    optimizer = optim.SGD(model.parameters(), lr=0.1)
    # 2. Define the scheduler (e.g., cut LR in half every 10 epochs)
    # https://docs.pytorch.org/docs/stable/optim.html
    scheduler = lr_scheduler.StepLR(optimizer, step_size=10, gamma=0.5)
    # 3. Training Loop
    for epoch in range(100):
        for inputs, targets in dataloader:
            optimizer.zero_grad()
            outputs = model(inputs)
            loss = criterion(outputs, targets)
            loss.backward()
            # Update model weights
            optimizer.step()
        # Update the learning rate AFTER the optimizer step at the end of the epoch
        scheduler.step()

    Regularization

    • Big models have a greater capacity to memorize
    • For relatively small datasets, this can be a problem
    • Weight decay prevents unchecked weight growth
      • Biases the network toward smaller weights (helpful)
      • Doesn’t prevent memorization
    • In deep models we use dropout
      • During training randomly disconnects neurons
      • Avoids solutions that are narrow and appropriate only for the training set
      • Forces the model to learn generalizable features spread across the model
      • Provably equivalent to learning an ensemble of models

    Dropout

    Diagram of a dropout layer inside a residual block: Linear, Dropout, tanh, Linear, Norm. During training, dropout draws a random 0 or 1 mask from a Bernoulli distribution with keep probability 1 minus p, multiplies the activations by it, and scales the survivors by 1 over 1 minus p, so each entry keeps the same expected value. A worked row with p = 0.5 shows half the values zeroed and the rest doubled, and a batch grid shows a fresh mask for every example. During evaluation there is no mask and no scaling. A fresh mask every training step means every step trains a different thinned network.

    Dropout destabilizes variance

    Helps with generalization when models are big and data is relatively small

    Pytorch

    • Dropout is a layer type in pytorch
    import torch
    import torch.nn as nn
    
    class ExampleNetwork(nn.Module):
        def __init__(self):
            super(ExampleNetwork, self):
            # 1. Define your layers
            self.fc1 = nn.Linear(784, 256)
            self.fc2 = nn.Linear(256, 10)
            # 2. Define the dropout layer (p = probability of zeroing an element)
            # Standard default is 0.5 if not specified
            self.dropout = nn.Dropout(p=0.3)
            self.relu = nn.ReLU()
    
        def forward(self, x):
            # 3. Pass the activation output through the dropout layer
            x = self.fc1(x)
            x = self.relu(x)
            x = self.dropout(x) # Applied after the activation function
            x = self.fc2(x)
            return x

    Project 1

    • Purpose of the project
      • Seeing and Applying course content to a practical problem
      • Getting used to searching for methods to improve performance and experimenting/innovating efficiently
        • Can’t run all of imagenet very fast, how to figure out quickly if your change is working?
      • Development good practice
        • The final models you develop must run with the provided harness script by simply pointing the harness at your repo without any mods to the script
    A colour photo of a flower split into its separate red, green and blue channel images.
    • 3 tasks
      • MNIST (easy)
        • Black and white hand written digit classification
        • 28x28 input with each being a grey scale pixel value
        • 60k images
        • 10 possible classes
      • Cifar-10 (harder)
        • RGB images of airplane, automobile, bird, cat, deer, dog, frog, horse, ship, truck
        • Each image is 3 32x32 matrices, one for each r, g, and b value for each pixel
        • 60k images
        • 10 classes
      • Imagenet 32x32 (bigger)
        • RGB images of a lot of different things
        • Each image is 3 32x32 matrices, one for each r, g, and b value for each pixel
        • 14million images
        • 1000 classes

    Understanding the base model

    We know much about this model.

    • For instance, what variance should the down projection weights have?
    • Why do we have normalization on each branch?
    • Why do we have a residual stream rather than FFN? Maybe we don’t need a big net for some tasks. Could FFN be better?
    • Why do we have an input projection?
    • Why do we have an output projection?

    Other things we don’t know.

    • What dimension should the residual stream have? Different across tasks?
    • What should the block latent dimensionality be? Task differences?
    • Does that activation function matter?
    • Would it help to scale the blocks by layer?
    • Do we need the first linear and tanh on the output? Is it helpful or not?
    • What is that softmax?
    Architecture diagram of the residual MLP base model. Input x (examples by input features) goes through a down projection into a 128-feature residual stream. Three residual blocks each apply Norm, Linear, tanh and Linear in their own hidden width and add the result back onto the stream. A final Norm feeds the output head: Linear, tanh, Linear to the number of classes, then softmax to give class probabilities.

    One hidden layer Multi-class Model

    import torch.nn as nn
    import torch
    class MLP_Multiclass(nn.Module):
        def __init__(self, features, hidden_size):
            super().__init__()
            self.fc1 = nn.Linear(features, hidden_size)
            self.fc2 = nn.Linear(hidden_size, 10)
        def forward(self, x):
            h = self.fc1(x)
            h = torch.tanh(h)
            return self.fc2(h)
        def predict(self, x):
            return torch.softmax(self.forward(x), dim=1).argmax(dim=1)
    Line chart of training and test loss per batch over about 9,000 training steps. Loss starts near 2.6, falls below 0.5 within the first few hundred batches, and flattens close to 0.05; the test loss tracks the training loss throughout, so there is no sign of overfitting.
    Forward pass doesn’t have softmax activation

    Predict (inference) does
    Confusion matrix of true digit against predicted digit for the 10 MNIST classes. Nearly all counts sit on the diagonal (between 866 and 1,120 correct per digit); the largest off-diagonal counts are 14 nines predicted as four, 11 zeros predicted as six, and 10 threes predicted as five.
    1 hidden layer
    20 epochs
    97.38% accurate
    Better than the base model…
    Bigger isn’t always better

    Project 1

    • Fork the starter repository
      • Has the project1 harness that will run the training and eval on the datasets
      • This is how I will compare your results across the class at the end
      • You are free to change this as you need to. For instance:
        • Only train and eval 1 task that you are working on improving
      • Keep in mind, I will be using the base version for final comparison
    • Improve the models
      • What works for 1 task may not be best for another, build different models for each task!
      • Apply the techniques we have discussed
        • Various activation functions and initializations
        • Change the latent size of the layer blocks
        • Change the dimension of the residual stream
        • Use different normalization layer configurations
        • Different optimizers
        • Regularization like dropout
        • Change the training method to have learning rate scheduling or data reuse
      • Experiment with things we haven’t talked about
        • Deep learning is large and there are many ideas out there
        • Find and try things that we haven’t discussed!
        • Build an entirely new architecture using linear layers, activations, normalization, residuals, dropout, etc.
      • Don’t use 1. any new layer types like convolutions or recurrence, 2. ensembles of multiple models, 3. pretrained weights, 4. extra data, 5. saved weights loaded in your training script
    • Be principled in your experimentation
      • Have a guess at what a change will do and document it with what actually happens through an experiment

    Project 1

    • Evaluate your imagenet model to understand how well it is functioning, required metrics are listed in canvas
      • Discuss the results in your report.
      • Is the gradient perfectly stable?
      • Is the variance stable?
      • Which branch is contributing most and which is contributing least?
      • Are the weights updating manageably?
      • Is the correlation low?
      • Is the activation function saturated?
    • Compare how your model had to change to get better performance on each task in the final report
    • Your grade will be based on
      • exploration of different ideas
      • documentation and analysis
      • your models working with the harness
      • your code being readable
    • Extra credit will go to the person whose results positively deviate most from the mean of the class
      • ReLU(Accuracy achieved on MNIST – class mean on MNIST)
      • + ReLU(Accuracy achieved on Cifar-10 – class mean on Cifar-10)
      • + ReLU(Accuracy achieved on imagenet– class mean on imagenet)
    • Be sure to point your eval notebook at your fork of the repository
      • Most aren’t currently