Understanding Linear Transformations in Deep Learning: Translation, Rotation, Scaling and ReLU Layers
Imagine you are playing with LEGO blocks on a table. You can slide a block to a new spot. You can spin it around. You can make it bigger or smaller.
That's exactly what a neural network does to numbers — slide, spin, stretch, and reshape them — until a picture of a cat turns into the word "cat".
In this tutorial we will understand every single step of that journey, starting from the most basic idea and building all the way up to how a Dense layer with ReLU activation works inside a real neural network.
Your Learning Road Map 🗺️
Here is what we will cover, step by step:
- 🟦 Translation — Sliding things to a new place
- 🔄 Rotation — Spinning things around a point
- 📐 Scaling — Making things bigger or smaller
- ➡️ Linear Transform — Rotation + Scaling combined, through the origin
- 🔀 Affine Transform — Linear transform + Translation combined
- 🧠 Dense Layer with ReLU — How a real neuron uses all of the above
🟡 Key insight before we start: Every single operation in a neural network — no matter how complex it looks — is built from these six simple ideas. Master these, and you understand the engine of all of AI.
Part 1 — Translation: Sliding Things Around 🟦
You draw a smiley face on a piece of paper. Now you pick it up and move it 3 centimetres to the right and 2 centimetres up. The face looks exactly the same — just in a new spot.
That's translation. Move something without changing its shape, size, or direction. Just shift its position.
In Numbers — What Does "Shift" Mean?
Every point in a picture (or a piece of data) has a position described by two numbers:
- x = how far left or right
- y = how far up or down
To translate (shift) a point, you simply add fixed amounts to x and y:
Original point: (x, y) Translation by: (tx, ty) ← the "shift amount" New point: (x + tx, y + ty)
That's the entire formula. Just addition!
Visual Diagram — Translation
Before translation After translating by (+3, +2) y y | | | ● (1, 2) | ● (4, 4) | | | | └──────── x └──────── x Point moved from (1,2) → (4,4) We added tx=3 to x and ty=2 to y
Python Code — Translation
Output:
Original point: [1 2] Translated point: [4 4]
See? Just adding two arrays. That's translation! 🎯
Translating Many Points At Once
Output:
Original square corners: [[0 0] [2 0] [2 2] [0 2]] Shifted square corners: [[5 3] [7 3] [7 5] [5 5]]
All four corners moved by the same amount. The square's shape is perfectly preserved — just in a new location. 🎉
🟢 Real-World Use: In image processing, translation moves an object within the frame. In data augmentation for training neural networks, images are randomly translated so the network learns that a cat in the corner is still a cat — same as a cat in the centre.
Part 2 — Rotation: Spinning Around a Point 🔄
Think of a clock hand. It spins around the centre of the clock. After 90 degrees, the 12 becomes the 3. After 180 degrees, 12 becomes 6.
Rotation spins every point in a shape around a fixed centre point, by a chosen angle. The shape stays the same size — it just faces a different direction.
The Rotation Formula
To rotate a point (x, y) by an angle θ (theta) around the origin (0, 0):
new_x = x × cos(θ) − y × sin(θ) new_y = x × sin(θ) + y × cos(θ)
This looks scary at first — but just think of it as a recipe. You put in (x, y) and the angle, and you get out the new (x, y). That's all it is!
Visual Diagram — Rotation by 90°
Before rotation After 90° counter-clockwise rotation y y | | | | ● (−2, 3) | ● (3, 2) | | | └──────── x └──────── x Original: (3, 2) After 90° CCW rotation: (−2, 3) (x and y swap, and the new x gets a minus sign)
Python Code — Rotation
Output:
Original: [3. 0.] After 90° rotation: [ 0. 3.]
The point moved from the x-axis (3, 0) to the y-axis (0, 3). A perfect 90-degree rotation! 🌀
Rotating a Full Shape
Output:
Original corners: [[1. 0.] [1. 1.] [0. 1.] [0. 0.]] Rotated corners (45°): [[ 0.7071 0.7071] [ 0. 1.4142] [-0.7071 0.7071] [ 0. 0. ]]
🟡 Important: The rotation formula always rotates around the origin (0, 0). If you want to rotate around a different point (like the centre of your shape), you must first translate that point to the origin, then rotate, then translate back. This is called "pivot rotation" and we will see it later!
🟢 Real-World Use: Rotation is used in data augmentation — randomly rotating training images helps the neural network learn that a rotated dog is still a dog. It's also the foundation of 3D graphics, robotics, and satellite navigation.
Part 3 — Scaling: Making Things Bigger or Smaller 📐
Imagine you have a toy car that is 10 cm long. You put it on a photocopier and set it to 200%. Now the copy is 20 cm long — twice as big.
Set it to 50% and the copy is 5 cm — half as big.
Scaling stretches or shrinks things by multiplying their coordinates by a number called the scale factor.
The Scaling Formula
Scale factor = s new_x = x × sx (sx = scale in x-direction) new_y = y × sy (sy = scale in y-direction) If sx = sy → uniform scaling (shape stays proportional) If sx ≠ sy → non-uniform scaling (shape gets stretched/squished)
Visual Diagram — Scaling
Original square Scaled by 2x (double size) Scaled by 0.5x (half size) (0,2)─────(2,2) (0,4)─────────(4,4) (0,1)──(1,1) │ │ │ │ │ │ │ │ │ │ │ │ (0,0)─────(2,0) (0,0)─────────(4,0) (0,0)──(1,0) Coordinates ×1 Coordinates ×2 Coordinates ×0.5
Python Code — Scaling
Output:
Original square: [[0. 0.] [2. 0.] [2. 2.] [0. 2.]] Scaled UP by 2x: [[0. 0.] [4. 0.] [4. 4.] [0. 4.]] Stretched (3x wide, 0.5x tall): [[0. 0.] [6. 0.] [6. 1.] [0. 1.]]
The uniform scale made everything twice as big. The non-uniform scale made it wide and flat — like a pancake! 🥞
🔴 DON'T confuse scaling with zooming a camera. Scaling actually changes the coordinate values of every point. Camera zoom is a viewing trick — it doesn't change the data. In neural networks, scaling changes the actual feature values, which directly affects what the network learns.
🟢 Real-World Use: Feature scaling (normalizing input data) is one of the most important preprocessing steps in machine learning. If one feature ranges from 0 to 1 and another from 0 to 1,000,000, the network trains much faster and more reliably when both are scaled to the same range — usually 0 to 1 or −1 to 1.
Part 4 — Linear Transform: Rotation + Scaling in One Step ➡️
So far we have done three separate moves: translate (slide), rotate (spin), and scale (resize).
A linear transform combines rotation and scaling (but NOT translation) into a single mathematical operation.
It's like saying: "Rotate the shape by 45° AND make it twice as big — in one go."
🟡 The Crucial Rule of Linear Transforms: A linear transform always passes through the origin (0, 0). This means the point (0, 0) never moves. If you need to move the origin, you need an affine transform — coming next!
The Linear Transform Formula
A linear transform is described by a matrix W. You apply it by multiplying every point by this matrix:
Output = W × Input
Where W is a 2×2 matrix (for 2D data):
W = [[w11, w12],
[w21, w22]]
For input point (x, y):
new_x = w11*x + w12*y
new_y = w21*x + w22*y
Special Cases — What Different W Matrices Do
Identity (do nothing):
W = [[1, 0], Input = (3, 2) Output = (3, 2)
[0, 1]]
Scale by 2:
W = [[2, 0], Input = (3, 2) Output = (6, 4)
[0, 2]]
Flip horizontally:
W = [[-1, 0], Input = (3, 2) Output = (-3, 2)
[ 0, 1]]
Shear (slant):
W = [[1, 0.5], Input = (2, 2) Output = (3, 2)
[0, 1 ]]
Python Code — Linear Transforms
Output:
Identity: [3. 2.] Scale x2: [6. 4.] Rotate 45°: [ 0.7071 3.5355] Shear: [4. 2.] Rotate + Scale: [ 1.4142 7.071 ]
Chaining Transforms — Order Matters!
Output:
Rotate then Scale: [0. 3.] Scale then Rotate: [0. 3.]
In this case the results are the same (rotation and uniform scaling commute). But for shears and non-uniform scales, order matters a lot! Always be explicit about which transform happens first.
🟢 This is the core of every neural network layer! The weight matrix W in a linear layer is exactly a linear transform. When your data passes through a layer, the network is rotating and scaling the data in high-dimensional space to separate different categories from each other.
Part 5 — Affine Transform: Linear + Translation = The Full Package 🔀
Remember: a linear transform can rotate and scale — but the origin never moves.
An affine transform adds one more power: it can also translate (slide the origin too).
Think of it like this:
- Linear = "Rotate and resize the stamp"
- Affine = "Rotate and resize the stamp, AND move it to a different spot on the page"
The Affine Transform Formula
Output = W × Input + b
Where:
W = the weight matrix (handles rotation + scaling)
b = the bias vector (handles translation / shifting)
For 2D input (x, y):
new_x = w11*x + w12*y + b1
new_y = w21*x + w22*y + b2
↑ ↑
Linear part Translation part
Visual Diagram — Affine Transform
Step 1: Apply linear transform (W × Input)
●(3,2) ──W──► ●(rotated+scaled position)
Step 2: Add bias (translate)
●(rotated+scaled) ──+b──► ●(final position)
Full formula: Output = W × Input + b
↑ ↑
Linear part Shift part
Python Code — Affine Transform
Output:
Input point: [2. 3.] Weight matrix W: [[ 1.4142 -1.4142] [ 1.4142 1.4142]] Bias vector b: [1. 5.] Affine output: [-0.4142 8.0711]
Applying Affine Transform to a Full Image Dataset
Output:
Input X (5 points): [[1. 2.] [3. 4.] [5. 6.] [7. 8.] [9. 0.]] Affine output (5 transformed points): [[-0.5 2.9] [-1.5 5.5] [-2.5 8.1] [-3.5 10.7] [ 5.5 3.7]]
All 5 data points transformed at once. This is exactly what happens when your data passes through a neural network layer! 🎉
🟡 This is the most important formula in all of deep learning:
Output = W × Input + b
Every dense (fully connected) layer in every neural network
performs this exact affine transform.
The weight matrix W and bias b are what the network learns during training.
Why Can't We Just Use Affine Transforms Alone? 🤔
Great question! Here's the problem:
If you stack 100 affine transforms one after another, the whole thing collapses down to just… one single affine transform. It's still just a straight line!
Real-world problems — like "is this a cat or a dog?" — can't be solved by a straight line. They need curves and bends.
Visual Proof — Linear Can't Solve XOR
XOR Problem (classic example):
Input (0,0) → Output: 0 Input (0,1) → Output: 1
Input (1,0) → Output: 1 Input (1,1) → Output: 0
On a grid:
y
|
1 | ○(0,1) ●(1,1)
| (XOR=1) (XOR=0)
0 | ●(0,0) ○(1,0)
| (XOR=0) (XOR=1)
└────────────────── x
0 1
● = output 0 ○ = output 1
Try to draw ONE straight line that separates ● from ○.
You can't! It's impossible without a curve.
This is exactly why neural networks have activation functions. They introduce the curves and bends that linear transforms cannot provide. The most popular activation function for modern networks is ReLU.
Part 6 — Dense Layer with ReLU Activation: Where It All Comes Together 🧠
Imagine a group of friends.
Each friend listens to all the information you have.
They each think about it in their own way (using their own W and b).
Then they make a decision: "Does this information excite me?"
If yes → they pass their excitement forward.
If no → they stay silent (send out a zero).
That "stay silent if not excited" part is exactly what ReLU does.
What is ReLU?
ReLU stands for Rectified Linear Unit. It has the simplest possible formula:
ReLU(x) = max(0, x) In plain English: If x is positive → keep it as-is If x is zero or negative → turn it to 0 Examples: ReLU(5.0) = 5.0 ✅ positive, keep it ReLU(0.3) = 0.3 ✅ positive, keep it ReLU(0.0) = 0.0 zero stays zero ReLU(-2.0) = 0.0 ❌ negative, turned to 0 ReLU(-99.9) = 0.0 ❌ negative, turned to 0
Visual Diagram — ReLU Function
Output │ ╱ 4 │ ╱ │ ╱ 3 │ ╱ │ ╱ 2 │ ╱ │ ╱ 1 │ ╱ │ ╱ 0 │╱─────────────────── Input -4 -3 -2 -1 0 1 2 3 4 Left of 0: flat line (everything = 0) Right of 0: straight diagonal (output = input) The "bend" at 0 is what makes ReLU non-linear!
A Complete Dense Layer — Step by Step
A Dense layer (also called Fully Connected layer) does two things in sequence:
- Step 1: Apply affine transform →
z = W × input + b - Step 2: Apply ReLU →
output = ReLU(z) = max(0, z)
┌─────────────────────────────────────────────────────────────┐ │ DENSE LAYER │ │ │ │ Input Affine Transform ReLU Activation │ │ [x1] │ │ [x2] ──► z = W @ x + b ──► output = max(0, z) │ │ [x3] │ │ │ │ W = learned weights b = learned bias │ └─────────────────────────────────────────────────────────────┘
Python Code — Dense Layer with ReLU from Scratch
Output:
Input X (3 samples × 4 features): [[ 0.8 0.3 -0.2 1.5] [ 0.1 -0.5 0.9 0.2] [ 1.2 0.7 0.4 -0.1]] Weights W (4 inputs → 3 neurons): [[ 0.2471 -0.2342 0.1041] [-0.4617 0.1826 0.2327] [ 0.0653 -0.0937 0.4979] [-0.1226 0.1765 -0.0872]] Bias b: [0. 0. 0.] Pre-activation z (before ReLU): [[ 0.0107 0.1889 0.0212] [-0.1132 -0.0259 0.5282] [ 0.0484 -0.1456 0.3714]] Output after ReLU (3 samples × 3 neurons): [[0.0107 0.1889 0.0212] [0. 0. 0.5282] [0.0484 0. 0.3714]]
Notice: some values in the pre-activation (z) were negative. After ReLU, those became exactly 0. The positive values passed through unchanged. 🎯
Putting It All Together — A 3-Layer Neural Network 🏗️
Now let's build a tiny but complete 3-layer neural network using everything we've learned. Each layer does: Affine transform → ReLU.
INPUT LAYER 1 LAYER 2 OUTPUT LAYER
(4 features) → (4→6 neurons) → (6→4 neurons) → (4→2 neurons)
W1, b1 W2, b2 W3, b3
+ ReLU + ReLU (no ReLU — raw scores)
Full Network — Python Code
Output:
=== Neural Network Forward Pass === After Layer 1 (4→6, ReLU): shape = (3, 6) [[ 0.4476 0. 0.3502 0. 0. 0. ] [ 0. 0.3258 0. 0.2801 0.0538 0. ] [ 0. 0. 0.5147 0. 0. 0.4219]] After Layer 2 (6→4, ReLU): shape = (3, 4) [[ 0.2111 0. 0.2973 0. ] [ 0. 0.1427 0. 0.3162] [ 0.3058 0. 0.4127 0. ]] Final Output (4→2, raw): shape = (3, 2) [[-0.1438 0.2541] [-0.0621 0.1873] [-0.1874 0.2924]] These are raw scores (logits) for 2 classes per sample.
Each of the 3 samples got two raw scores — one per class. The class with the higher score becomes the network's prediction!
🟢 This is exactly what GPT, image classifiers, and recommendation systems do — just with millions of neurons instead of 6. The fundamental operations are identical to what you just ran above. You have built the mathematical engine of modern AI from scratch! 🏆
Why ReLU? Why Not Other Activations? 🤷
Before ReLU became popular (around 2012), the go-to activations were sigmoid and tanh. Here's why ReLU won:
| Property | Sigmoid / Tanh | ReLU |
|---|---|---|
| Computation speed | Slow (needs exp() function) | ✅ Extremely fast (just compare to 0) |
| Vanishing gradient | ❌ Severe — gradients shrink to near zero | ✅ No vanishing for positive values |
| Sparsity | All neurons always active | ✅ Only some neurons fire — more efficient |
| Deep networks | ❌ Hard to train beyond 3–4 layers | ✅ Works well for 100+ layers |
| Dying neurons | Rare | ⚠️ Possible if learning rate too high |
🟡 Dying ReLU: If many neurons receive large negative values constantly, they get "stuck" at zero and never activate again — they are "dead". The fix is Leaky ReLU: instead of outputting 0 for negative inputs, it outputs a tiny fraction (like 0.01 × input). This keeps all neurons alive even for very negative values.
Output:
Input: [-3. -1. 0. 1. 3.] ReLU: [0. 0. 0. 1. 3.] Leaky ReLU: [-0.03 -0.01 0. 1. 3. ] ELU: [-0.9502 -0.6321 0. 1. 3. ]
Building the Same Network in Keras — 5 Lines! ⚡
Everything above is what happens inside the box when you write this in Keras (TensorFlow's high-level API):
Output:
Model: "sequential" _________________________________________________________________ Layer (type) Output Shape Param # ================================================================= dense (Dense) (None, 6) 30 ← 4×6 weights + 6 biases dense_1 (Dense) (None, 4) 28 ← 6×4 weights + 4 biases dense_2 (Dense) (None, 2) 10 ← 4×2 weights + 2 biases ================================================================= Total params: 68 (272.00 Byte) _________________________________________________________________
Same structure — 3 layers, same sizes — just expressed in 5 clean lines instead of raw NumPy. Keras handles weight initialization, batching, and backpropagation for you.
🟢 DO:
Use Keras for real projects — it's fast, battle-tested, and handles GPU acceleration.
But now you understand exactly what Keras is doing behind every Dense layer!
output = relu(W @ input + b) — that's it, every time.
Beginner Mistakes — And How to Fix Them 🚨
Mistake 1: Forgetting to Scale Input Data
Output:
Normalized input: [[0. 1.] [1. 0.]]
Mistake 2: Using ReLU in the Output Layer
🔴 DON'T use ReLU on the final output layer of a classifier. ReLU kills negative values — but those negative raw scores (logits) carry important information about which class is less likely. Remove them and your network learns garbage.
Mistake 3: Zero-Initializing All Weights
Mistake 4: Making the Network Too Deep Too Fast
🟡 Golden Rule: Start with the simplest possible network. If training accuracy is low → add more neurons or layers. If validation accuracy is much lower than training → add dropout or reduce size. Never start by building the biggest network you can imagine.
Real-World Applications — Where These Live 🌍
- Translation in AI: Data augmentation — shifting training images randomly prevents the network from "memorizing" exact pixel positions.
- Rotation in AI: Rotating training images teaches robustness. Also the core of how 3D game engines and robots calculate positions.
- Scaling in AI: Feature normalization is a must-do preprocessing step. Without it, training is unstable and slow. MinMaxScaler and StandardScaler in sklearn implement this.
- Linear Transform in AI: The weight matrix W in every Dense layer is literally a linear transform applied to the input data.
- Affine Transform in AI: Dense layer = affine transform. Convolutional layer = local affine transform with shared weights. Both are everywhere in modern AI.
- Dense + ReLU in AI: Used in image classifiers, language models, recommendation engines, fraud detection, medical diagnosis, autonomous driving — essentially every AI application in production today.
Quick Summary 📝
What we learned today:
-
Translation
→ Add a fixed vector:
new = point + t→ slides without changing shape - Rotation → Multiply by a rotation matrix R → spins around the origin, preserves distances
- Scaling → Multiply by a diagonal matrix S → resizes uniformly or non-uniformly
- Linear Transform → Multiply by any matrix W → rotate + scale + shear, always through origin
-
Affine Transform
→
output = W @ input + b→ linear transform + translation → this is the Dense layer formula! -
Dense + ReLU
→
output = relu(W @ input + b)→ affine transform + non-linearity → the building block of every neural network
🟢 The Master Formula of Deep Learning:
output = relu( W @ input + b )
Translation lives in b (bias).
Rotation and Scaling live in W (weight matrix).
Non-linearity lives in relu.
Stack this 100 times — you have a deep neural network.
That's the whole engine. You now understand it completely. 🏆
Keep building and keep experimenting! Every complex neural network in existence — GPT-5, image generators, voice assistants — is built from these same simple blocks you just mastered. Happy learning! 🐼✨
Comments
Post a Comment