Skip to main content

CNN Fundamentals: Convolution, Filters, Padding, Stride, and Pooling

Calculating read time…

A convolutional neural network learns to recognize patterns in grid-shaped data — most commonly images — by sliding small, learnable filters across the input instead of connecting every input pixel to every neuron. Convolution, filters, padding, stride, and pooling are the five mechanical pieces that make this work: one decides what each filter computes, one decides what it looks for, and the other three decide exactly how the sliding window moves and how the resulting feature maps shrink as they pass deeper into the network. 🖼️

This matters in production because every one of these five choices directly changes a layer's output shape, its parameter count, and its receptive field — get the arithmetic wrong and a network either crashes with a shape error or silently trains on far less spatial context than intended. Engineers who can compute an output shape by hand, rather than trusting a framework's default, debug architecture mismatches in minutes instead of hours. ⚙️

A diagram of a 5 by 5 input grid with a one-cell zero-padding border shown in a lighter shade around a 3 by 3 core, a small 3 by 3 filter highlighted over one 3 by 3 region of the input, an arrow to a multiply-and-sum operation, and an arrow into a single highlighted cell of a smaller output feature map grid, illustrating that one filter position produces exactly one output value.

Figure 1. Original diagram: the same small filter is reused at every position across the padded input.

🔀 Quick Comparison: The Five CNN Building Blocks

Concept What it controls Has learnable parameters? Typical effect on output size
Convolution How each output value is computed from a local input region Yes (the filter weights and bias) Depends on kernel size, padding, and stride together
Filter (kernel) What visual pattern a channel detects Yes, learned via backpropagation Kernel size shrinks output unless padding compensates
Padding How the input's borders are extended before sliding the filter No Increases output size
Stride How many positions the filter jumps between steps No Decreases output size as stride grows
Pooling Summarizing a small region into one value No Decreases output size, usually by the pooling window's stride

1. Foundations: Why Convolution Instead of a Fully Connected Layer

🧠 Child-friendly analogy first: imagine looking for your cat in a large photo by holding up a small picture-frame cutout the exact size of "a cat's face" and sliding it across the whole photo, checking each spot. You don't need a separate, custom-shaped frame for every single position in the photo — the same small frame works everywhere, because a cat's face looks like a cat's face no matter where in the photo it appears. That reusable, sliding frame is exactly what a convolutional filter is.

A fully connected layer, of the kind covered in this series' perceptron and forward-pass articles, would connect every single input pixel to every output unit independently — for a modest 224-by-224 color image, that's already over 150,000 input values feeding into every unit, with a completely separate set of weights learned for every possible position in the image. Convolution instead reuses one small set of weights (the filter) across every spatial position, which is both dramatically cheaper in parameters and, because the same pattern-detector is applied everywhere, naturally able to recognize a pattern regardless of where in the image it appears.

PyTorch's own documentation for its 2D convolution layer defines the operation directly as a sum, for every output channel, of the cross-correlation between that channel's learned weight and each input channel, plus a learnable bias term. (PyTorch Conv2d documentation) Notably, the documentation is explicit that this is the cross-correlation operator rather than the flipped-kernel operation that pure mathematics usually calls "convolution" — a naming carryover from signal processing that doesn't change how the layer is trained or used in practice.

🎯 Use this when explaining to a teammate why CNNs need far fewer parameters than an equivalent fully connected network for the same image size.

2. Filters: What a Convolution Actually Learns

🧠 Analogy: think of a filter as a small stencil with numbers instead of holes. Laying that stencil over a patch of the image, multiplying each stencil number by the pixel value underneath it, and adding everything together produces one number that's large when the underlying patch resembles the pattern the stencil numbers encode, and small or negative when it doesn't.

What it does: a filter (also called a kernel) is a small grid of learnable numbers — commonly 3×3 or 5×5 — that gets multiplied elementwise against every local patch of the input it slides over, with the results summed into a single output value plus a bias, exactly as PyTorch's Conv2d documentation defines it.

Why it's needed: a single filter alone can only detect one kind of local pattern; a convolutional layer typically applies many independent filters to the same input simultaneously, each producing its own output channel, so that different filters can specialize in detecting different patterns — edges at different orientations, color contrasts, or simple textures in early layers, with deeper layers building on those into more complex shapes.

What fails without understanding it: a common source of confusion is expecting a single filter to somehow detect "everything" a layer needs; in reality, the number of output channels (equivalently, the number of independent filters in that layer) is a specific architectural choice that directly trades off representational capacity against parameter count and compute.

🎯 Use this when deciding how many output channels to give a new convolutional layer.

3. Padding: Controlling What Happens at the Edges

🧠 Analogy: imagine that picture-frame cutout again, but now think about what happens as it reaches the very edge of the photo — part of the frame would hang off into empty space. Padding is the decision to either stop the frame right at the photo's true edge (losing a little coverage near the border) or to first tape a blank border around the photo so the frame can still be centered fully over every original pixel, including the ones near the edge.

What it does: padding adds extra values, most commonly zeros, around the border of the input before the filter slides over it. PyTorch's Conv2d documentation allows padding to be specified as an explicit integer or tuple of pixels, or as the strings "valid" (no padding at all) or "same" (padding chosen automatically so the output spatial size matches the input, for stride 1). (PyTorch Conv2d documentation)

Why it's needed: without padding, every convolution shrinks its spatial dimensions somewhat, because the filter can only be centered where it fully fits inside the input. Stacking many convolutional layers with no padding compounds this shrinkage, which can either be a deliberate design choice or an accidental one that leaves too little spatial resolution for later layers to work with.

What fails without it: pixels near the border of an unpadded input are covered by fewer filter positions than pixels near the center, meaning edge information systematically contributes less to the output than center information — a subtle bias that padding specifically corrects for, at the cost of processing some entirely artificial zero values along the border.

🎯 Use this when deciding whether a convolutional stack should preserve spatial resolution layer to layer or deliberately shrink it.

4. Stride: How Far the Filter Moves Each Step

🧠 Analogy: picture that same picture-frame search again, but now imagine sliding the frame one full frame-width at a time instead of pixel by pixel — you'll cover the whole photo in far fewer moves, at the cost of never checking the in-between positions at all. That's exactly the trade-off stride makes.

What it does: stride is the number of positions the filter shifts between one output value and the next, in each spatial direction. PyTorch's Conv2d documentation defines stride as controlling the step of the cross-correlation, and defaults it to 1, meaning the filter moves one position at a time by default. (PyTorch Conv2d documentation)

Why it's needed: a stride greater than 1 shrinks the output map's spatial size in roughly direct proportion to the stride, which reduces the compute and memory needed for every subsequent layer — an alternative to pooling for deliberately downsampling a feature map, sometimes used directly inside a convolutional layer instead of as a separate pooling step.

What fails without understanding it: because output size depends jointly on input size, kernel size, padding, and stride, changing any one of these four values without recomputing the resulting shape is one of the most common sources of an unexpected shape mismatch the very first time a new architecture is run.

🎯 Use this when you need to shrink a feature map's spatial size directly inside a convolutional layer rather than with a separate pooling layer.

5. Pooling: Shrinking Feature Maps on Purpose

🧠 Analogy: imagine summarizing a paragraph by keeping only its single most important sentence and discarding the rest — you lose some detail, but you keep the strongest signal and end up with something much shorter to work with. Max pooling does exactly this to small regions of a feature map: keep only the strongest activation in each region, discard the rest.

What it does: PyTorch's documentation for its 2D max pooling layer defines the output at each position as the maximum value found within a small sliding window over the input, with that window's size, stride, and padding configured independently of any convolutional layer nearby; the default stride equals the window's own kernel size, so non-overlapping windows are the default behavior. (PyTorch MaxPool2d documentation)

A diagram of a 4 by 4 feature map divided into four non-overlapping 2 by 2 regions, each shaded a different color, with an arrow showing each region's four values collapsing to a single value equal to the maximum of that region, producing a 2 by 2 output map that is one quarter the size of the input.

Figure 2. Original diagram: a 2×2 max pooling window with stride 2 keeps only the strongest value in each region.

Why it's needed: pooling reduces the spatial size of feature maps with no learnable parameters at all, cutting compute and memory for every following layer, while also providing a small amount of translation tolerance — a pattern shifted by a pixel or two often still produces the same maximum value within a pooling window.

What fails without it: PyTorch's own MaxPool2d documentation notes that when padding is used with max pooling, the input is implicitly padded with negative infinity rather than zero, specifically so padded positions can never be mistakenly selected as the maximum. Using ordinary zero padding logic when reasoning about max pooling by hand is a common source of manual calculation errors.

🎯 Use this when deliberately downsampling a feature map between convolutional stages while adding zero extra parameters.

6. Real Example: LeNet-5 and Gradient-Based Document Recognition

Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner's 1998 paper "Gradient-Based Learning Applied to Document Recognition" describes LeNet-5, one of the earliest convolutional architectures trained end-to-end with backpropagation, applying alternating convolutional and subsampling (pooling) layers to handwritten character recognition. The paper's own abstract states directly that convolutional neural networks, specifically designed to deal with the variability of two-dimensional shapes, are shown to outperform all other techniques compared on a standard handwritten digit recognition task. (LeCun, Bottou, Bengio & Haffner, "Gradient-Based Learning Applied to Document Recognition," Proceedings of the IEEE, 1998)

✅ Worked example: the same paper's abstract describes a deployed application built on this approach: a graph transformer network using convolutional character recognizers for reading bank cheques, stated to be deployed commercially and reading several million cheques per day at the time of publication — a documented, narrowly scoped real-world deployment of the exact convolution-plus-pooling pattern this article covers.

This is one specific, citable data point about one architecture's documented performance and one documented deployment, not a claim that every later CNN architecture shares that exact result; the broader adoption of CNNs across computer vision reflects a much larger body of subsequent research beyond the scope of this article.

🎯 Use this when you want a concrete, citable historical anchor for why convolution plus pooling became a foundational pattern in computer vision.

7. Implementation Walkthrough

The short, illustrative examples below show the concepts above as code. They are simplified teaching snippets, not production-ready modules.

# Illustrative example: computing a convolution's output size by hand
def conv_output_size(input_size, kernel_size, padding, stride):
    return (input_size + 2 * padding - kernel_size) // stride + 1

# a 32x32 input, 3x3 kernel, padding=1, stride=1 -> stays 32x32
print(conv_output_size(32, 3, 1, 1))   # -> 32

# the same input with no padding shrinks to 30x30
print(conv_output_size(32, 3, 0, 1))   # -> 30
# Illustrative example: a small conv-pool stack, PyTorch-style
model = build_sequential(
    conv2d_layer(in_channels=3, out_channels=16, kernel_size=3, padding=1),
    relu(),
    max_pool2d_layer(kernel_size=2, stride=2),   # halves height and width
    conv2d_layer(in_channels=16, out_channels=32, kernel_size=3, padding=1),
    relu(),
    max_pool2d_layer(kernel_size=2, stride=2),   # halves again
)

batch = create_tensor(shape=(8, 3, 64, 64))
output = model.forward(batch)
print(output.shape)   # -> (8, 32, 16, 16)

Best practices: compute and log the expected output shape after every convolutional or pooling layer when designing a new architecture, rather than only discovering a mismatch when the model actually runs, and be explicit about padding rather than relying on remembering a framework's default.

🎯 Use this when sizing a new convolutional architecture or debugging a shape mismatch between two stacked layers.

8. Enterprise Rollout: Shape Contracts for Vision Pipelines

Convolutional architectures have more interacting shape-determining parameters than a simple feedforward network, which makes explicit shape governance especially valuable once a vision model moves toward production.

Ownership and governance: document the expected input resolution, channel count, and every layer's kernel size, padding, and stride for a production vision model, since these jointly determine the model's receptive field — how much of the original image each late-layer output value is actually influenced by — which is a property worth tracking deliberately rather than leaving implicit.

CI gates: add a test that runs a fixed, known input resolution through the full architecture and asserts the exact output shape at each stage, so a change to any single layer's padding or stride that shifts downstream shapes is caught immediately rather than surfacing as a mismatch several layers later.

Dataset and preprocessing versioning: pin the exact input resolution and resizing or cropping strategy used to produce training data, since a CNN's learned filters are tuned to a specific expected input scale, and silently changing preprocessing between training and serving can degrade accuracy without any explicit error.

Dashboards and alerts: track inference latency and memory per image resolution actually seen in production, since an unexpectedly large input resolution reaching a model can quietly multiply the compute cost of every convolutional layer far more than a similar-looking increase would for a simple fully connected layer.

Incident response: when a vision model produces a shape error or unexpectedly poor accuracy after a deployment, capture the exact input resolution and channel count that triggered it, since a mismatch between the resolution a model was trained on and the resolution it's actually receiving in production is one of the most common, and most quickly diagnosable, root causes.

🎯 Use this when a computer vision model built in research is being handed off to a team that will serve it at scale.

9. Common Mistakes

Assuming a convolutional layer preserves spatial size by default. PyTorch's Conv2d defaults padding to 0, so an unpadded convolution shrinks its output relative to its input unless padding is explicitly set. The production impact is a stack of several convolutional layers producing a much smaller final feature map than the architecture designer intended, often only discovered when a later layer's expected input size doesn't match.

Confusing max pooling's negative-infinity padding with ordinary zero padding. PyTorch's MaxPool2d documentation explicitly states padding for this layer uses negative infinity, not zero, specifically so padded positions can never be selected as the maximum. Reasoning about a padded max-pooling operation using zero-padding logic produces an incorrect output value at the borders.

Not accounting for stride and pooling compounding across many layers. Because each stride-2 convolution or 2×2 pooling layer roughly halves spatial resolution, stacking several of them in sequence shrinks the feature map exponentially, not linearly. The production impact is a final feature map that's far smaller than expected, sometimes shrinking to a size too small for a later layer's kernel to even fit.

Treating "convolution" in deep learning frameworks as identical to the flipped-kernel operation from signal processing. PyTorch's own documentation clarifies its Conv2d implements cross-correlation, not the flipped-kernel convolution operator from classical signal processing. This distinction rarely matters for training a network from scratch, since the filter is learned either way, but it can matter when porting hand-designed, fixed filter weights from a signal-processing context.

Choosing an arbitrary number of filters per layer without a clear rationale. The number of output channels in each convolutional layer is a direct, deliberate trade-off between representational capacity, parameter count, and compute cost; picking it arbitrarily, rather than considering how it interacts with the rest of the architecture, is a common source of either underpowered or needlessly expensive models.

❓ FAQ

What's the actual difference between stride and pooling, if both shrink the output?

A convolution with stride greater than 1 computes fewer output positions using its own learnable filter weights, while pooling is a separate, parameter-free layer that summarizes small regions (typically by taking the maximum) with no weights to learn at all. Both shrink spatial size, but only one of them has anything to train.

Why does PyTorch call Conv2d a "cross-correlation" rather than a true convolution?

PyTorch's own documentation states directly that the star operator in its formula denotes the valid 2D cross-correlation operator, which skips the kernel-flipping step that classical mathematical convolution includes; since the filter's weights are learned rather than fixed, this naming difference doesn't change how the network is trained.

Does "same" padding always keep the output exactly the same size as the input?

PyTorch's Conv2d documentation offers "same" as a padding option specifically for keeping the output spatial size equal to the input's, but this guarantee applies for stride 1; combining "same"-style padding with a stride greater than 1 changes the output size regardless of the padding chosen, since stride itself skips output positions.

Why do CNN filters share the same weights at every spatial position?

Weight sharing is what lets one filter detect the same pattern no matter where it appears in the image, and it's also what keeps the parameter count of a convolutional layer independent of the input's spatial size, unlike a fully connected layer where every input position needs its own dedicated weights.

Is LeNet-5 still used in production today?

LeNet-5 itself, as described in the 1998 paper, is primarily of historical and educational interest today rather than a common choice for new production systems; the article's real-example section cites it specifically for its documented historical results and deployment, not as a current architectural recommendation.

🔗 References & Further Reading

📝 Summary

  • Convolution reuses one small, learnable filter across every spatial position, dramatically cutting parameters compared to a fully connected layer.
  • Filters are what a layer actually learns to detect; the number of output channels controls how many independent patterns a layer can look for at once.
  • Padding controls what happens at the input's edges and can keep output size equal to input size at stride 1.
  • Stride controls how far the filter jumps between steps, directly shrinking output size as it grows.
  • Pooling is a parameter-free way to shrink feature maps, and PyTorch's max pooling specifically pads with negative infinity to avoid ever selecting a padded position as the maximum.
  • LeNet-5's original 1998 paper documents both a benchmark result and a real commercial deployment of this exact convolution-plus-pooling pattern.

From a sliding picture frame to a paragraph boiled down to its best sentence — that's convolution, padding, stride, and pooling, working together. Happy building! 🚀

Comments