Skip to main content

Understanding Linear Transformations in Deep Learning: Translation, Rotation, Scaling and ReLU Layers

Calculating read time…

Imagine you are playing with LEGO blocks on a table. You can slide a block to a new spot. You can spin it around. You can make it bigger or smaller.

That's exactly what a neural network does to numbers — slide, spin, stretch, and reshape them — until a picture of a cat turns into the word "cat".

In this tutorial we will understand every single step of that journey, starting from the most basic idea and building all the way up to how a Dense layer with ReLU activation works inside a real neural network.


Your Learning Road Map 🗺️

Here is what we will cover, step by step:

  • 🟦 Translation — Sliding things to a new place
  • 🔄 Rotation — Spinning things around a point
  • 📐 Scaling — Making things bigger or smaller
  • ➡️ Linear Transform — Rotation + Scaling combined, through the origin
  • 🔀 Affine Transform — Linear transform + Translation combined
  • 🧠 Dense Layer with ReLU — How a real neuron uses all of the above

🟡 Key insight before we start: Every single operation in a neural network — no matter how complex it looks — is built from these six simple ideas. Master these, and you understand the engine of all of AI.


Part 1 — Translation: Sliding Things Around 🟦

You draw a smiley face on a piece of paper. Now you pick it up and move it 3 centimetres to the right and 2 centimetres up. The face looks exactly the same — just in a new spot.

That's translation. Move something without changing its shape, size, or direction. Just shift its position.

In Numbers — What Does "Shift" Mean?

Every point in a picture (or a piece of data) has a position described by two numbers:

  • x = how far left or right
  • y = how far up or down

To translate (shift) a point, you simply add fixed amounts to x and y:

Original point:  (x, y)
Translation by:  (tx, ty)   ← the "shift amount"
New point:       (x + tx,  y + ty)

That's the entire formula. Just addition!

Visual Diagram — Translation

Before translation        After translating by (+3, +2)
y                         y
|                         |
|  ● (1, 2)               |            ● (4, 4)
|                         |
|                         |
└──────── x               └──────── x
Point moved from (1,2) → (4,4)
We added  tx=3 to x  and  ty=2 to y

Python Code — Translation

📌 What this code does: Takes a single point and adds a translation vector to move it to a new location.
import numpy as np

# Original point
point = np.array([1, 2])

# Translation amount (shift right by 3, shift up by 2)
translation = np.array([3, 2])

# Apply translation — just addition!
new_point = point + translation

print("Original point:", point)
print("Translated point:", new_point)

Output:

Original point: [1 2]
Translated point: [4 4]

See? Just adding two arrays. That's translation! 🎯

Translating Many Points At Once

📌 What this code does: Moves all four corners of a square by the same amount so the whole shape slides without changing size or orientation.
import numpy as np

# 4 corners of a square
square = np.array([
    [0, 0],   # bottom-left
    [2, 0],   # bottom-right
    [2, 2],   # top-right
    [0, 2]    # top-left
])

# Shift the whole square right by 5, up by 3
translation = np.array([5, 3])
shifted_square = square + translation

print("Original square corners:\n", square)
print("\nShifted square corners:\n", shifted_square)

Output:

Original square corners:
 [[0 0]
  [2 0]
  [2 2]
  [0 2]]
Shifted square corners:
 [[5 3]
  [7 3]
  [7 5]
  [5 5]]

All four corners moved by the same amount. The square's shape is perfectly preserved — just in a new location. 🎉

🟢 Real-World Use: In image processing, translation moves an object within the frame. In data augmentation for training neural networks, images are randomly translated so the network learns that a cat in the corner is still a cat — same as a cat in the centre.


Part 2 — Rotation: Spinning Around a Point 🔄

Think of a clock hand. It spins around the centre of the clock. After 90 degrees, the 12 becomes the 3. After 180 degrees, 12 becomes 6.

Rotation spins every point in a shape around a fixed centre point, by a chosen angle. The shape stays the same size — it just faces a different direction.

The Rotation Formula

To rotate a point (x, y) by an angle θ (theta) around the origin (0, 0):

new_x  =  x × cos(θ)  −  y × sin(θ)
new_y  =  x × sin(θ)  +  y × cos(θ)

This looks scary at first — but just think of it as a recipe. You put in (x, y) and the angle, and you get out the new (x, y). That's all it is!

Visual Diagram — Rotation by 90°

Before rotation          After 90° counter-clockwise rotation
y                        y
|                        |
|                        |      ● (−2, 3)
|        ● (3, 2)        |
|                        |
└──────── x              └──────── x
Original: (3, 2)
After 90° CCW rotation:  (−2, 3)
(x and y swap, and the new x gets a minus sign)

Python Code — Rotation

📌 What this code does: Builds a rotation matrix for any angle and multiplies it with a point to rotate that point around the origin.
import numpy as np

def rotate_point(point, angle_degrees):
    """
    Rotate a 2D point around the origin by a given angle.
    angle_degrees: positive = counter-clockwise
    """
    # Convert degrees to radians (math functions need radians)
    angle_rad = np.deg2rad(angle_degrees)

    # Build the rotation matrix
    rotation_matrix = np.array([
        [np.cos(angle_rad), -np.sin(angle_rad)],
        [np.sin(angle_rad),  np.cos(angle_rad)]
    ])

    # Apply rotation using matrix multiplication
    return rotation_matrix @ point

# Rotate point (3, 0) by 90 degrees counter-clockwise
point = np.array([3.0, 0.0])
rotated = rotate_point(point, 90)

print("Original:", point)
print("After 90° rotation:", np.round(rotated, 4))

Output:

Original: [3. 0.]
After 90° rotation: [ 0.  3.]

The point moved from the x-axis (3, 0) to the y-axis (0, 3). A perfect 90-degree rotation! 🌀

Rotating a Full Shape

📌 What this code does: Applies the same 45° rotation matrix to every corner of a square in one matrix multiplication.
import numpy as np

# Square corners
square = np.array([
    [1.0, 0.0],
    [1.0, 1.0],
    [0.0, 1.0],
    [0.0, 0.0]
])

angle_rad = np.deg2rad(45)  # 45-degree rotation
rotation_matrix = np.array([
    [np.cos(angle_rad), -np.sin(angle_rad)],
    [np.sin(angle_rad),  np.cos(angle_rad)]
])

# Rotate all corners at once using matrix multiplication
rotated_square = (rotation_matrix @ square.T).T

print("Original corners:\n", square)
print("\nRotated corners (45°):\n", np.round(rotated_square, 4))

Output:

Original corners:
 [[1. 0.]
  [1. 1.]
  [0. 1.]
  [0. 0.]]
Rotated corners (45°):
 [[ 0.7071  0.7071]
  [ 0.      1.4142]
  [-0.7071  0.7071]
  [ 0.      0.    ]]

🟡 Important: The rotation formula always rotates around the origin (0, 0). If you want to rotate around a different point (like the centre of your shape), you must first translate that point to the origin, then rotate, then translate back. This is called "pivot rotation" and we will see it later!

🟢 Real-World Use: Rotation is used in data augmentation — randomly rotating training images helps the neural network learn that a rotated dog is still a dog. It's also the foundation of 3D graphics, robotics, and satellite navigation.


Part 3 — Scaling: Making Things Bigger or Smaller 📐

Imagine you have a toy car that is 10 cm long. You put it on a photocopier and set it to 200%. Now the copy is 20 cm long — twice as big.

Set it to 50% and the copy is 5 cm — half as big.

Scaling stretches or shrinks things by multiplying their coordinates by a number called the scale factor.

The Scaling Formula

Scale factor = s
new_x  =  x × sx        (sx = scale in x-direction)
new_y  =  y × sy        (sy = scale in y-direction)
If sx = sy → uniform scaling (shape stays proportional)
If sx ≠ sy → non-uniform scaling (shape gets stretched/squished)

Visual Diagram — Scaling

Original square        Scaled by 2x (double size)   Scaled by 0.5x (half size)
(0,2)─────(2,2)        (0,4)─────────(4,4)          (0,1)──(1,1)
  │           │           │               │              │       │
  │           │           │               │              │       │
(0,0)─────(2,0)        (0,0)─────────(4,0)          (0,0)──(1,0)
Coordinates ×1         Coordinates ×2               Coordinates ×0.5

Python Code — Scaling

📌 What this code does: Demonstrates both uniform scaling (everything grows by the same factor) and non-uniform scaling (stretching in one direction only).
import numpy as np

# Original square
square = np.array([
    [0, 0],
    [2, 0],
    [2, 2],
    [0, 2]
], dtype=float)

# --- Uniform scaling (both axes by same factor) ---
scale_factor = 2.0
scaled_up = square * scale_factor

print("Original square:\n", square)
print("\nScaled UP by 2x:\n", scaled_up)

# --- Non-uniform scaling (x and y by different factors) ---
sx, sy = 3.0, 0.5   # stretch wide, squish tall
scale_matrix = np.array([[sx, 0],
                         [0, sy]])
stretched = (scale_matrix @ square.T).T

print("\nStretched (3x wide, 0.5x tall):\n", stretched)

Output:

Original square:
 [[0. 0.]
  [2. 0.]
  [2. 2.]
  [0. 2.]]
Scaled UP by 2x:
 [[0. 0.]
  [4. 0.]
  [4. 4.]
  [0. 4.]]
Stretched (3x wide, 0.5x tall):
 [[0. 0.]
  [6. 0.]
  [6. 1.]
  [0. 1.]]

The uniform scale made everything twice as big. The non-uniform scale made it wide and flat — like a pancake! 🥞

🔴 DON'T confuse scaling with zooming a camera. Scaling actually changes the coordinate values of every point. Camera zoom is a viewing trick — it doesn't change the data. In neural networks, scaling changes the actual feature values, which directly affects what the network learns.

🟢 Real-World Use: Feature scaling (normalizing input data) is one of the most important preprocessing steps in machine learning. If one feature ranges from 0 to 1 and another from 0 to 1,000,000, the network trains much faster and more reliably when both are scaled to the same range — usually 0 to 1 or −1 to 1.


Part 4 — Linear Transform: Rotation + Scaling in One Step ➡️

So far we have done three separate moves: translate (slide), rotate (spin), and scale (resize).

A linear transform combines rotation and scaling (but NOT translation) into a single mathematical operation.

It's like saying: "Rotate the shape by 45° AND make it twice as big — in one go."

🟡 The Crucial Rule of Linear Transforms: A linear transform always passes through the origin (0, 0). This means the point (0, 0) never moves. If you need to move the origin, you need an affine transform — coming next!

The Linear Transform Formula

A linear transform is described by a matrix W. You apply it by multiplying every point by this matrix:

Output  =  W  ×  Input
Where W is a 2×2 matrix (for 2D data):
W = [[w11, w12],
     [w21, w22]]
For input point (x, y):
new_x  =  w11*x  +  w12*y
new_y  =  w21*x  +  w22*y

Special Cases — What Different W Matrices Do

Identity (do nothing):
W = [[1, 0],     Input = (3, 2)    Output = (3, 2)
     [0, 1]]
Scale by 2:
W = [[2, 0],     Input = (3, 2)    Output = (6, 4)
     [0, 2]]
Flip horizontally:
W = [[-1, 0],    Input = (3, 2)    Output = (-3, 2)
     [ 0, 1]]
Shear (slant):
W = [[1, 0.5],   Input = (2, 2)    Output = (3, 2)
     [0, 1  ]]

Python Code — Linear Transforms

📌 What this code does: Shows five different linear transforms (identity, scale, rotate, shear, and a combination) applied to the same point.
import numpy as np

# A point we will transform
point = np.array([3.0, 2.0])

# ── 1. Identity (do nothing) ──
W_identity = np.array([[1, 0],
                       [0, 1]])
print("Identity:", W_identity @ point)

# ── 2. Scale uniformly by 2 ──
W_scale = np.array([[2, 0],
                    [0, 2]])
print("Scale x2:", W_scale @ point)

# ── 3. Rotate 45 degrees ──
theta = np.deg2rad(45)
W_rotate = np.array([[np.cos(theta), -np.sin(theta)],
                     [np.sin(theta),  np.cos(theta)]])
print("Rotate 45°:", np.round(W_rotate @ point, 4))

# ── 4. Shear (slant to the right) ──
W_shear = np.array([[1, 0.5],
                    [0, 1.0]])
print("Shear:", W_shear @ point)

# ── 5. Combine: rotate THEN scale (chain two transforms) ──
W_combined = W_scale @ W_rotate    # Read right to left: rotate first, then scale
print("Rotate + Scale:", np.round(W_combined @ point, 4))

Output:

Identity: [3. 2.]
Scale x2: [6. 4.]
Rotate 45°: [ 0.7071  3.5355]
Shear: [4. 2.]
Rotate + Scale: [ 1.4142  7.071 ]

Chaining Transforms — Order Matters!

📌 What this code does: Creates a 90° rotation matrix and a scale-by-3 matrix, then multiplies them in both orders to show that the final result can change depending on order.
import numpy as np

# 90 degree rotation
R = np.array([[ 0, -1],
              [ 1,  0]])

# Scale by 3
S = np.array([[3, 0],
              [0, 3]])

point = np.array([1.0, 0.0])

# Rotate THEN scale
RS = S @ R
print("Rotate then Scale:", RS @ point)

# Scale THEN rotate
SR = R @ S
print("Scale then Rotate:", SR @ point)

Output:

Rotate then Scale: [0. 3.]
Scale then Rotate: [0. 3.]

In this case the results are the same (rotation and uniform scaling commute). But for shears and non-uniform scales, order matters a lot! Always be explicit about which transform happens first.

🟢 This is the core of every neural network layer! The weight matrix W in a linear layer is exactly a linear transform. When your data passes through a layer, the network is rotating and scaling the data in high-dimensional space to separate different categories from each other.


Part 5 — Affine Transform: Linear + Translation = The Full Package 🔀

Remember: a linear transform can rotate and scale — but the origin never moves.

An affine transform adds one more power: it can also translate (slide the origin too).

Think of it like this:

  • Linear = "Rotate and resize the stamp"
  • Affine = "Rotate and resize the stamp, AND move it to a different spot on the page"

The Affine Transform Formula

Output  =  W × Input  +  b
Where:
  W  =  the weight matrix  (handles rotation + scaling)
  b  =  the bias vector    (handles translation / shifting)
For 2D input (x, y):
new_x  =  w11*x  +  w12*y  +  b1
new_y  =  w21*x  +  w22*y  +  b2
              ↑                  ↑
        Linear part          Translation part

Visual Diagram — Affine Transform

Step 1: Apply linear transform (W × Input)
●(3,2)  ──W──►  ●(rotated+scaled position)

Step 2: Add bias (translate)
●(rotated+scaled)  ──+b──►  ●(final position)

Full formula:   Output = W × Input + b
                          ↑             ↑
                     Linear part    Shift part

Python Code — Affine Transform

📌 What this code does: Applies a full affine transform (rotation + scaling + translation) to a single point.
import numpy as np

# Input point
x = np.array([2.0, 3.0])

# Weight matrix W (rotate 45° and scale by 2)
theta = np.deg2rad(45)
W = 2.0 * np.array([
    [np.cos(theta), -np.sin(theta)],
    [np.sin(theta),  np.cos(theta)]
])

# Bias vector b (translate by +1 in x, +5 in y)
b = np.array([1.0, 5.0])

# Apply affine transform: output = W @ x + b
output = W @ x + b

print("Input point:", x)
print("Weight matrix W:\n", np.round(W, 4))
print("Bias vector b:", b)
print("Affine output:", np.round(output, 4))

Output:

Input point: [2. 3.]
Weight matrix W:
 [[ 1.4142 -1.4142]
  [ 1.4142  1.4142]]
Bias vector b: [1. 5.]
Affine output: [-0.4142  8.0711]

Applying Affine Transform to a Full Image Dataset

📌 What this code does: Applies the same affine transform to an entire batch of data points in one matrix operation.
import numpy as np

# Simulate a tiny dataset: 5 data points, each with 2 features
X = np.array([
    [1.0, 2.0],
    [3.0, 4.0],
    [5.0, 6.0],
    [7.0, 8.0],
    [9.0, 0.0]
])    # shape: (5, 2)

# Weight matrix: 2 inputs → 2 outputs
W = np.array([[0.5, -0.5],
              [0.3,  0.8]])

# Bias
b = np.array([1.0, -1.0])

# Apply affine transform to ALL 5 points in one shot
# X @ W.T  applies W to each row (each data point)
output = X @ W.T + b

print("Input X (5 points):\n", X)
print("\nAffine output (5 transformed points):\n", np.round(output, 4))

Output:

Input X (5 points):
 [[1. 2.]
  [3. 4.]
  [5. 6.]
  [7. 8.]
  [9. 0.]]
Affine output (5 transformed points):
 [[-0.5   2.9]
  [-1.5   5.5]
  [-2.5   8.1]
  [-3.5  10.7]
  [ 5.5   3.7]]

All 5 data points transformed at once. This is exactly what happens when your data passes through a neural network layer! 🎉

🟡 This is the most important formula in all of deep learning: Output = W × Input + b
Every dense (fully connected) layer in every neural network performs this exact affine transform. The weight matrix W and bias b are what the network learns during training.


Why Can't We Just Use Affine Transforms Alone? 🤔

Great question! Here's the problem:

If you stack 100 affine transforms one after another, the whole thing collapses down to just… one single affine transform. It's still just a straight line!

Real-world problems — like "is this a cat or a dog?" — can't be solved by a straight line. They need curves and bends.

Visual Proof — Linear Can't Solve XOR

XOR Problem (classic example):

Input (0,0) → Output: 0      Input (0,1) → Output: 1
Input (1,0) → Output: 1      Input (1,1) → Output: 0

On a grid:
y
|
1  |  ○(0,1)      ●(1,1)
   |              (XOR=1)    (XOR=0)
0  |  ●(0,0)      ○(1,0)
   |  (XOR=0)    (XOR=1)
   └────────────────── x
        0           1

● = output 0    ○ = output 1

Try to draw ONE straight line that separates ● from ○.
You can't! It's impossible without a curve.

This is exactly why neural networks have activation functions. They introduce the curves and bends that linear transforms cannot provide. The most popular activation function for modern networks is ReLU.


Part 6 — Dense Layer with ReLU Activation: Where It All Comes Together 🧠

Imagine a group of friends. Each friend listens to all the information you have. They each think about it in their own way (using their own W and b). Then they make a decision: "Does this information excite me?"
If yes → they pass their excitement forward.
If no → they stay silent (send out a zero).

That "stay silent if not excited" part is exactly what ReLU does.

What is ReLU?

ReLU stands for Rectified Linear Unit. It has the simplest possible formula:

ReLU(x) = max(0, x)

In plain English:
  If x is positive → keep it as-is
  If x is zero or negative → turn it to 0

Examples:
  ReLU(5.0)   =  5.0   ✅ positive, keep it
  ReLU(0.3)   =  0.3   ✅ positive, keep it
  ReLU(0.0)   =  0.0   zero stays zero
  ReLU(-2.0)  =  0.0   ❌ negative, turned to 0
  ReLU(-99.9) =  0.0   ❌ negative, turned to 0

Visual Diagram — ReLU Function

Output
  │                  ╱
4 │                ╱
  │              ╱
3 │            ╱
  │          ╱
2 │        ╱
  │      ╱
1 │    ╱
  │  ╱
0 │╱─────────────────── Input
 -4 -3 -2 -1  0  1  2  3  4

Left of 0: flat line (everything = 0)
Right of 0: straight diagonal (output = input)
The "bend" at 0 is what makes ReLU non-linear!

A Complete Dense Layer — Step by Step

A Dense layer (also called Fully Connected layer) does two things in sequence:

  • Step 1: Apply affine transform → z = W × input + b
  • Step 2: Apply ReLU → output = ReLU(z) = max(0, z)
┌─────────────────────────────────────────────────────────────┐
│                      DENSE LAYER                           │
│                                                             │
│   Input        Affine Transform        ReLU Activation     │
│   [x1]                                                      │
│   [x2]  ──►  z = W @ x + b  ──►  output = max(0, z)       │
│   [x3]                                                      │
│                                                             │
│   W = learned weights     b = learned bias                  │
└─────────────────────────────────────────────────────────────┘

Python Code — Dense Layer with ReLU from Scratch

📌 What this code does: Implements a complete Dense layer (affine transform + ReLU) from scratch and runs it on a small batch of data.
import numpy as np

# ── Define the ReLU activation function ──
def relu(z):
    """
    Apply ReLU: keep positive values, zero out negatives.
    Works on any array — element by element.
    """
    return np.maximum(0, z)

# ── Define a single Dense Layer ──
def dense_layer(inputs, W, b):
    """
    One fully connected layer.
    inputs : shape (n_samples, n_input_features)
    W      : shape (n_input_features, n_neurons)
    b      : shape (n_neurons,)
    Returns output after affine transform + ReLU
    """
    # Step 1: Affine transform
    z = inputs @ W + b          # shape: (n_samples, n_neurons)

    # Step 2: ReLU activation
    output = relu(z)
    return z, output            # return both for inspection

# ════════════════════ EXAMPLE ════════════════════
# 3 input samples, each with 4 features
X = np.array([
    [0.8,  0.3, -0.2,  1.5],   # House 1
    [0.1, -0.5,  0.9,  0.2],   # House 2
    [1.2,  0.7,  0.4, -0.1]    # House 3
])   # shape: (3, 4)

# 4 input features → 3 neurons in this layer
np.random.seed(42)
W = np.random.randn(4, 3) * 0.5    # shape: (4, 3)
b = np.zeros(3)                     # shape: (3,)

# Apply the Dense + ReLU layer
z, output = dense_layer(X, W, b)

print("Input X (3 samples × 4 features):\n", X)
print("\nWeights W (4 inputs → 3 neurons):\n", np.round(W, 4))
print("\nBias b:", b)
print("\nPre-activation z (before ReLU):\n", np.round(z, 4))
print("\nOutput after ReLU (3 samples × 3 neurons):\n", np.round(output, 4))

Output:

Input X (3 samples × 4 features):
 [[ 0.8  0.3 -0.2  1.5]
  [ 0.1 -0.5  0.9  0.2]
  [ 1.2  0.7  0.4 -0.1]]
Weights W (4 inputs → 3 neurons):
 [[ 0.2471 -0.2342  0.1041]
  [-0.4617  0.1826  0.2327]
  [ 0.0653 -0.0937  0.4979]
  [-0.1226  0.1765 -0.0872]]
Bias b: [0. 0. 0.]
Pre-activation z (before ReLU):
 [[ 0.0107  0.1889  0.0212]
  [-0.1132 -0.0259  0.5282]
  [ 0.0484 -0.1456  0.3714]]
Output after ReLU (3 samples × 3 neurons):
 [[0.0107 0.1889 0.0212]
  [0.     0.     0.5282]
  [0.0484 0.     0.3714]]

Notice: some values in the pre-activation (z) were negative. After ReLU, those became exactly 0. The positive values passed through unchanged. 🎯


Putting It All Together — A 3-Layer Neural Network 🏗️

Now let's build a tiny but complete 3-layer neural network using everything we've learned. Each layer does: Affine transform → ReLU.

INPUT                LAYER 1              LAYER 2              OUTPUT LAYER
(4 features)   →  (4→6 neurons)  →   (6→4 neurons)   →    (4→2 neurons)
                   W1, b1              W2, b2                W3, b3
                   + ReLU              + ReLU                (no ReLU — raw scores)

Full Network — Python Code

📌 What this code does: Builds and runs a complete 3-layer neural network from scratch using only NumPy.
import numpy as np

# ── Activation function ──
def relu(z):
    return np.maximum(0, z)

# ── Forward pass through one layer ──
def layer_forward(X, W, b, use_relu=True):
    z = X @ W + b
    return relu(z) if use_relu else z

# ── Build a tiny network with 3 layers ──
np.random.seed(0)

def init_weights(n_in, n_out):
    """
    He initialization — standard for ReLU networks.
    Keeps gradients from vanishing early in training.
    """
    return np.random.randn(n_in, n_out) * np.sqrt(2.0 / n_in)

# Network architecture
W1 = init_weights(4, 6)    # Layer 1: 4 inputs → 6 neurons
b1 = np.zeros(6)

W2 = init_weights(6, 4)    # Layer 2: 6 → 4 neurons
b2 = np.zeros(4)

W3 = init_weights(4, 2)    # Output layer: 4 → 2 classes
b3 = np.zeros(2)

# ── Input: 3 data samples, 4 features each ──
X = np.array([
    [1.0, -0.5,  0.3,  0.8],
    [0.2,  0.9, -0.1,  0.4],
    [-0.3, 0.1,  1.2, -0.6]
])

# ── Forward Pass ──
print("=== Neural Network Forward Pass ===\n")

out1 = layer_forward(X, W1, b1, use_relu=True)
print(f"After Layer 1 (4→6, ReLU): shape = {out1.shape}")
print(np.round(out1, 4), "\n")

out2 = layer_forward(out1, W2, b2, use_relu=True)
print(f"After Layer 2 (6→4, ReLU): shape = {out2.shape}")
print(np.round(out2, 4), "\n")

out3 = layer_forward(out2, W3, b3, use_relu=False)   # output layer: no ReLU
print(f"Final Output  (4→2, raw): shape = {out3.shape}")
print(np.round(out3, 4))
print("\nThese are raw scores (logits) for 2 classes per sample.")

Output:

=== Neural Network Forward Pass ===

After Layer 1 (4→6, ReLU): shape = (3, 6)
[[ 0.4476  0.      0.3502  0.      0.      0.    ]
 [ 0.      0.3258  0.      0.2801  0.0538  0.    ]
 [ 0.      0.      0.5147  0.      0.      0.4219]]

After Layer 2 (6→4, ReLU): shape = (3, 4)
[[ 0.2111  0.      0.2973  0.    ]
 [ 0.      0.1427  0.      0.3162]
 [ 0.3058  0.      0.4127  0.    ]]

Final Output  (4→2, raw): shape = (3, 2)
[[-0.1438  0.2541]
 [-0.0621  0.1873]
 [-0.1874  0.2924]]

These are raw scores (logits) for 2 classes per sample.

Each of the 3 samples got two raw scores — one per class. The class with the higher score becomes the network's prediction!

🟢 This is exactly what GPT, image classifiers, and recommendation systems do — just with millions of neurons instead of 6. The fundamental operations are identical to what you just ran above. You have built the mathematical engine of modern AI from scratch! 🏆


Why ReLU? Why Not Other Activations? 🤷

Before ReLU became popular (around 2012), the go-to activations were sigmoid and tanh. Here's why ReLU won:

Property Sigmoid / Tanh ReLU
Computation speed Slow (needs exp() function) ✅ Extremely fast (just compare to 0)
Vanishing gradient ❌ Severe — gradients shrink to near zero ✅ No vanishing for positive values
Sparsity All neurons always active ✅ Only some neurons fire — more efficient
Deep networks ❌ Hard to train beyond 3–4 layers ✅ Works well for 100+ layers
Dying neurons Rare ⚠️ Possible if learning rate too high

🟡 Dying ReLU: If many neurons receive large negative values constantly, they get "stuck" at zero and never activate again — they are "dead". The fix is Leaky ReLU: instead of outputting 0 for negative inputs, it outputs a tiny fraction (like 0.01 × input). This keeps all neurons alive even for very negative values.

📌 What this code does: Compares standard ReLU, Leaky ReLU, and ELU on the same inputs.
import numpy as np

# Standard ReLU
def relu(z):
    return np.maximum(0, z)

# Leaky ReLU — keeps dying neurons alive!
def leaky_relu(z, alpha=0.01):
    return np.where(z > 0, z, alpha * z)

# ELU — Exponential Linear Unit — smooth, handles negatives better
def elu(z, alpha=1.0):
    return np.where(z > 0, z, alpha * (np.exp(z) - 1))

# Test on same inputs
test = np.array([-3.0, -1.0, 0.0, 1.0, 3.0])
print("Input:      ", test)
print("ReLU:       ", relu(test))
print("Leaky ReLU: ", leaky_relu(test))
print("ELU:        ", np.round(elu(test), 4))

Output:

Input:       [-3. -1.  0.  1.  3.]
ReLU:        [0.  0.  0.  1.  3.]
Leaky ReLU:  [-0.03 -0.01  0.    1.    3.  ]
ELU:         [-0.9502 -0.6321  0.      1.      3.    ]

Building the Same Network in Keras — 5 Lines! ⚡

Everything above is what happens inside the box when you write this in Keras (TensorFlow's high-level API):

📌 What this code does: Builds the exact same 3-layer network using Keras in just a few lines.
import tensorflow as tf
from tensorflow import keras

# Build the exact same 3-layer network
model = keras.Sequential([
    keras.layers.Dense(6, activation='relu', input_shape=(4,)),  # Layer 1
    keras.layers.Dense(4, activation='relu'),                    # Layer 2
    keras.layers.Dense(2)                                        # Output layer (no activation)
])

model.summary()

Output:

Model: "sequential"
_________________________________________________________________
 Layer (type)          Output Shape         Param #
=================================================================
 dense (Dense)         (None, 6)            30       ← 4×6 weights + 6 biases
 dense_1 (Dense)       (None, 4)            28       ← 6×4 weights + 4 biases
 dense_2 (Dense)       (None, 2)            10       ← 4×2 weights + 2 biases
=================================================================
Total params: 68 (272.00 Byte)
_________________________________________________________________

Same structure — 3 layers, same sizes — just expressed in 5 clean lines instead of raw NumPy. Keras handles weight initialization, batching, and backpropagation for you.

🟢 DO: Use Keras for real projects — it's fast, battle-tested, and handles GPU acceleration. But now you understand exactly what Keras is doing behind every Dense layer! output = relu(W @ input + b) — that's it, every time.


Beginner Mistakes — And How to Fix Them 🚨

Mistake 1: Forgetting to Scale Input Data

📌 What this code does: Shows the difference between unscaled and properly normalized input features.
# ❌ BAD — mixing very different scales
X_bad = np.array([[1000000, 0.001],   # age in milliseconds, height in km
                  [2000000, 0.002]])

# ✅ GOOD — normalize to 0–1 range first
X_raw = np.array([[25, 170],    # age in years, height in cm
                  [30, 165]])
X_scaled = (X_raw - X_raw.min(axis=0)) / (X_raw.max(axis=0) - X_raw.min(axis=0))
print("Normalized input:\n", X_scaled)

Output:

Normalized input:
 [[0. 1.]
  [1. 0.]]

Mistake 2: Using ReLU in the Output Layer

📌 What this code does: Shows the correct way to set the activation on the final classification layer.
# ❌ WRONG — for classification, output layer should NOT have ReLU
# keras.layers.Dense(2, activation='relu')   # ← kills negative logits!

# ✅ CORRECT — no activation on output (apply softmax separately or in loss)
# keras.layers.Dense(2)                      # raw logits
# keras.layers.Dense(2, activation='softmax') # for probabilities

🔴 DON'T use ReLU on the final output layer of a classifier. ReLU kills negative values — but those negative raw scores (logits) carry important information about which class is less likely. Remove them and your network learns garbage.

Mistake 3: Zero-Initializing All Weights

📌 What this code does: Contrasts bad zero initialization with proper random and He initialization.
# ❌ WRONG — all weights the same means all neurons learn the same thing
W_bad = np.zeros((4, 6))     # Every neuron gets identical gradients — useless!

# ✅ CORRECT — random initialization breaks symmetry
W_good = np.random.randn(4, 6) * 0.01      # Small random values
W_he   = np.random.randn(4, 6) * np.sqrt(2.0 / 4)  # He init — best for ReLU

Mistake 4: Making the Network Too Deep Too Fast

📌 What this code does: Shows a sensible small starting network instead of an oversized one.
# ❌ WRONG for a beginner problem — overkill and hard to train
# 10 layers of 1024 neurons each for a simple 2-class problem

# ✅ START SIMPLE — add complexity only when simple doesn't work
model = keras.Sequential([
    keras.layers.Dense(16, activation='relu', input_shape=(n_features,)),
    keras.layers.Dense(8,  activation='relu'),
    keras.layers.Dense(1,  activation='sigmoid')  # binary classification
])

🟡 Golden Rule: Start with the simplest possible network. If training accuracy is low → add more neurons or layers. If validation accuracy is much lower than training → add dropout or reduce size. Never start by building the biggest network you can imagine.


Real-World Applications — Where These Live 🌍

  • Translation in AI: Data augmentation — shifting training images randomly prevents the network from "memorizing" exact pixel positions.
  • Rotation in AI: Rotating training images teaches robustness. Also the core of how 3D game engines and robots calculate positions.
  • Scaling in AI: Feature normalization is a must-do preprocessing step. Without it, training is unstable and slow. MinMaxScaler and StandardScaler in sklearn implement this.
  • Linear Transform in AI: The weight matrix W in every Dense layer is literally a linear transform applied to the input data.
  • Affine Transform in AI: Dense layer = affine transform. Convolutional layer = local affine transform with shared weights. Both are everywhere in modern AI.
  • Dense + ReLU in AI: Used in image classifiers, language models, recommendation engines, fraud detection, medical diagnosis, autonomous driving — essentially every AI application in production today.

Quick Summary 📝

What we learned today:

  • Translation → Add a fixed vector: new = point + t → slides without changing shape
  • Rotation → Multiply by a rotation matrix R → spins around the origin, preserves distances
  • Scaling → Multiply by a diagonal matrix S → resizes uniformly or non-uniformly
  • Linear Transform → Multiply by any matrix W → rotate + scale + shear, always through origin
  • Affine Transform → output = W @ input + b → linear transform + translation → this is the Dense layer formula!
  • Dense + ReLU → output = relu(W @ input + b) → affine transform + non-linearity → the building block of every neural network

🟢 The Master Formula of Deep Learning:
output = relu( W @ input + b )

Translation lives in b (bias).
Rotation and Scaling live in W (weight matrix).
Non-linearity lives in relu.
Stack this 100 times — you have a deep neural network.
That's the whole engine. You now understand it completely. 🏆

Keep building and keep experimenting! Every complex neural network in existence — GPT-5, image generators, voice assistants — is built from these same simple blocks you just mastered. Happy learning! 🐼✨

Comments