TRENDING
Rows of identical brass-colored apartment mailboxes with small locks and name labels along an orange corridor wall
October 9, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
Street-level upward view of the Monetary Authority of Singapore building and neighbouring office towers under a pale sky
October 9, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
Cast-iron late Qing dynasty coin minting press with a large flywheel, displayed in a museum case
October 9, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google
Rows of closed oak library card catalog drawers, each with a brass pull and a blank label holder
October 9, 2026
How to Encrypt PII in Python and Keep It Searchable With Blind Indexes
Close-up of a vintage Western Electric manual telephone switchboard with orange lamps, red patch cords plugged into jacks, a rotary dial and a black handset
October 9, 2026
Microsoft’s Agent Lightning v1.0 Turns Agent Training Into a Sample-Accounting Problem
09 Oct 2026
SXZ.io SXZ.io
  • Home
Search the Site
Popular Searches:
Technology Amazon AI
Recent Posts
Two orange safety relief valves on grey pressure vessels in an industrial plant
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
Yellow diamond-shaped merging traffic warning sign showing a side road joining a main road
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A lugworm lying on wet sand and mud at low tide
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
SXZ.io SXZ.io
  • Home

Categories

Articles 232 Posts
News 234 Posts
Learning Hub 204 Posts
Home/Learning Hub/How to Build a Neural Network From Scratch in Python With NumPy
Learning Hub

How to Build a Neural Network From Scratch in Python With NumPy

A hands-on, from-scratch walkthrough of forward propagation, backpropagation, and gradient descent in plain NumPy: build a single neuron, watch it fail on a problem it cannot solve, fix it with a...

August 23, 2026 17 Min Read
45

Every time an app recognizes your handwriting, flags a spam email, or recommends a video, there is very likely a neural network doing the work behind the scenes. It is tempting to assume that means the code underneath must be enormously complicated: advanced calculus, a wall of matrix notation, thousands of lines you could never fully follow. In reality, the core mechanism behind almost every neural network you have heard of, from a simple spam filter to a large language model, comes down to a small set of repeated ideas: multiply some numbers by weights, add them up, squash the result through a simple function, measure how wrong the answer was, and nudge the numbers slightly so the next answer is a little less wrong. Repeat that a few thousand times and the network learns.

Table Of Content

  • Core Concepts, Defined Before We Use Them
  • Prerequisites
  • Step 1: Set Up Your Environment
  • Step 2: Build and Train a Single Neuron
  • Step 3: Watch the Same Neuron Fail, and Understand Exactly Why
  • Step 4: Add a Hidden Layer
  • Common Mistake: Initializing All Weights to Zero
  • Common Mistake: Picking a Learning Rate Blindly
  • Step 5: Compare With a Real Framework (PyTorch)
  • How to Confirm Everything Worked End to End
  • A Note on the Loss Function, and What to Learn Next

This tutorial builds a working neural network completely from scratch in Python, using nothing but NumPy for the math. No machine learning framework, no black box. You will write the forward pass, the loss calculation, and the backpropagation update rule yourself, line by line, so that when you later reach for a framework like PyTorch or TensorFlow, you understand what it is automating instead of treating it as magic. Along the way you will watch a real, deliberately too-simple network fail on a problem it cannot solve, understand precisely why it fails, fix it by adding a hidden layer, and verify with real captured output (not textbook numbers) that the fix works. The tutorial closes by rebuilding the same network in PyTorch in about a dozen lines, so you can see exactly what a framework does for you under the hood.

Core Concepts, Defined Before We Use Them

  • Neuron: the smallest unit of a neural network. It takes in one or more numbers, multiplies each by a weight, adds a bias, and passes the result through an activation function.
  • Weight: a number that controls how strongly one input influences a neuron’s output. Weights are what the network actually learns during training.
  • Bias: an extra learned number added to a neuron’s weighted sum, letting it shift its decision threshold independently of the inputs.
  • Activation function: a simple function, usually non-linear, applied to a neuron’s weighted sum. This tutorial uses the sigmoid function throughout. Without a non-linear activation function, stacking multiple layers would still only be able to represent a straight line; Step 4 shows exactly why.
  • Forward propagation: running input data through the network, layer by layer, to produce a prediction.
  • Loss function: a single number that measures how wrong the network’s prediction was compared to the correct answer. This tutorial uses mean squared error.
  • Gradient descent: the optimization method that adjusts each weight and bias by a small step in the direction that reduces the loss.
  • Backpropagation: the algorithm that efficiently computes exactly how much each individual weight contributed to the loss, using the chain rule from calculus, so gradient descent knows which direction to nudge each one.
  • Epoch: one complete pass of training over the entire dataset.

Prerequisites

  • Python 3.10 or later. This tutorial was built and tested on Python 3.13.14 on Windows; the code is plain, portable Python and NumPy, and behaves identically on Linux or macOS.
  • pip, to install NumPy 2.5.2 (the version used here) and, optionally, PyTorch for the comparison in Step 5.
  • Comfortable reading basic Python: functions, loops, lists, and NumPy arrays. No prior machine learning experience is assumed.
  • Basic algebra (multiplying and adding numbers). Some light calculus (the chain rule) shows up in backpropagation, but it is explained in plain terms as it comes up, not assumed going in.

Step 1: Set Up Your Environment

Create a folder for this project, set up a virtual environment inside it, and install NumPy:

mkdir nn-from-scratch
cd nn-from-scratch
python -m venv venv
venv\Scripts\activate
pip install numpy

On Linux or macOS, activate the virtual environment with source venv/bin/activate instead of the Windows-specific command above. NumPy’s own documentation covers the array operations this tutorial relies on (@ for matrix multiplication, broadcasting, and elementwise math) in far more depth than this tutorial needs. Once the virtual environment is active, confirm NumPy installed correctly:

python -c "import numpy; print(numpy.__version__)"

Expected output (a slightly newer version number is fine if you install later):

2.5.2

Step 2: Build and Train a Single Neuron

Start with the simplest possible network: one neuron, two inputs. Picture a small monitoring script watching two independent failed-login checks on a server, check_a and check_b. You want to flag the event for the audit log if at least one check trips:

  • Neither check trips (0, 0): nothing to flag, target is 0.
  • Only check_a trips (1, 0): flag it, target is 1.
  • Only check_b trips (0, 1): flag it, target is 1.
  • Both trip (1, 1): still flag it, target is 1.

This is the logical OR function, and it has a property mathematicians call linear separability: if you plot those four (check_a, check_b) points on a 2D grid, you can draw a single straight line that puts the “0” case on one side and every “1” case on the other. Hold onto that property; it matters a lot in Step 3.

Here is the neuron’s math before any code. Given inputs x, weights w, and a bias b, the neuron first computes a weighted sum z = (x . w) + b, then squashes that through the sigmoid activation function, a = 1 / (1 + e^-z), which maps any real number into the range 0 to 1 so its output can be read as a probability.

Save this as single_neuron.py:

import numpy as np

np.random.seed(42)


def sigmoid(z):
    return 1 / (1 + np.exp(-z))


# Two independent failed-login checks on a server. Flag for the audit log
# if AT LEAST ONE of them trips.
X = np.array([[0, 0], [0, 1], [1, 0], [1, 1]], dtype=float)
y = np.array([[0], [1], [1], [1]], dtype=float)

w = np.random.randn(2, 1) * 0.1
b = np.zeros((1, 1))

lr = 0.5
epochs = 5000

for epoch in range(epochs):
    z = X @ w + b
    a = sigmoid(z)

    loss = np.mean((a - y) ** 2)

    d_loss_d_a = 2 * (a - y) / len(X)
    d_a_d_z = a * (1 - a)
    delta = d_loss_d_a * d_a_d_z

    dw = X.T @ delta
    db = np.sum(delta, axis=0, keepdims=True)

    w -= lr * dw
    b -= lr * db

    if epoch % 1000 == 0:
        print(f"epoch {epoch:5d}  loss {loss:.4f}")

print(f"epoch {epochs:5d}  loss {loss:.4f}")
print()
print("check_a  check_b  target  predicted  rounded")
z = X @ w + b
a = sigmoid(z)
for xi, yi, ai in zip(X, y, a):
    print(f"   {int(xi[0])}        {int(xi[1])}       {int(yi[0])}     {ai[0]:.4f}      {round(ai[0])}")

Run it:

python single_neuron.py

Expected output:

epoch     0  loss 0.2456
epoch  1000  loss 0.0064
epoch  2000  loss 0.0029
epoch  3000  loss 0.0018
epoch  4000  loss 0.0013
epoch  5000  loss 0.0011

check_a  check_b  target  predicted  rounded
   0        0       0     0.0487      0
   0        1       1     0.9695      1
   1        0       1     0.9695      1
   1        1       1     0.9999      1

For every one of the 5,000 epochs, the code did four things: ran the inputs through the neuron (forward propagation), measured how far off the prediction was (the loss), used the chain rule to work out exactly how much nudging the weights and bias would reduce that loss (backpropagation, even though with a single neuron there is nothing to propagate back through yet), and took a small step in that direction (gradient descent). The loss drops from about 0.25 to about 0.001, and every prediction now rounds to the correct target. The single neuron solved OR because OR is linearly separable and a single neuron, geometrically, is only capable of drawing one straight decision boundary.

Step 3: Watch the Same Neuron Fail, and Understand Exactly Why

Now try a problem that looks almost identical, but is not. Picture two redundant overheat sensors bolted to the same server rack, installed specifically so one can catch what the other misses. If both are silent, nothing is happening. If both trip together, that is a real event, and a separate automated system already handles it. What is actually worth paging a human about is the case where the two sensors disagree: one tripped and the other did not, which usually means a sensor fault rather than an actual overheating rack.

  • Neither sensor trips (0, 0): normal, target is 0.
  • Only sensor_a trips (1, 0): disagreement, flag it, target is 1.
  • Only sensor_b trips (0, 1): disagreement, flag it, target is 1.
  • Both trip (1, 1): a real event, already handled elsewhere, target is 0.

This is the logical XOR (exclusive or) function. Try plotting those same four points again: (0, 0) and (1, 1) need to land on one side of a dividing line, while (1, 0) and (0, 1) need to land on the other. Try to draw one straight line that manages that. You cannot; the two “1” points sit on opposite corners from each other, and so do the two “0” points. XOR is the textbook example of a function that is not linearly separable, and a single neuron with a sigmoid activation function can only ever draw one straight decision boundary, no matter how long you train it or how carefully you tune the learning rate. This is not a new observation: Marvin Minsky and Seymour Papert’s 1969 book Perceptrons formally proved that a single-layer perceptron cannot solve any linearly nonseparable problem, including XOR, a result that briefly slowed neural network research until multi-layer networks (what Step 4 builds next) were shown to solve it.

Run the identical code from Step 2, changing only the four target values from OR’s [0, 1, 1, 1] to XOR’s [0, 1, 1, 0]. Save it as single_neuron_fails.py:

import numpy as np

np.random.seed(42)


def sigmoid(z):
    return 1 / (1 + np.exp(-z))


# Two redundant overheat sensors on the same rack. Flag for manual
# inspection only when they DISAGREE (one tripped, the other did not),
# since that usually means a sensor fault rather than a real event.
X = np.array([[0, 0], [0, 1], [1, 0], [1, 1]], dtype=float)
y = np.array([[0], [1], [1], [0]], dtype=float)

w = np.random.randn(2, 1) * 0.1
b = np.zeros((1, 1))

lr = 0.5
epochs = 5000

for epoch in range(epochs):
    z = X @ w + b
    a = sigmoid(z)

    loss = np.mean((a - y) ** 2)

    d_loss_d_a = 2 * (a - y) / len(X)
    d_a_d_z = a * (1 - a)
    delta = d_loss_d_a * d_a_d_z

    dw = X.T @ delta
    db = np.sum(delta, axis=0, keepdims=True)

    w -= lr * dw
    b -= lr * db

    if epoch % 1000 == 0:
        print(f"epoch {epoch:5d}  loss {loss:.4f}")

print(f"epoch {epochs:5d}  loss {loss:.4f}")
print()
print("sensor_a  sensor_b  target  predicted  rounded  result")
z = X @ w + b
a = sigmoid(z)
n_correct = 0
for xi, yi, ai in zip(X, y, a):
    pred = round(ai[0])
    ok = pred == yi[0]
    n_correct += ok
    print(f"    {int(xi[0])}        {int(xi[1])}       {int(yi[0])}     {ai[0]:.4f}      {pred}      {'OK' if ok else 'WRONG'}")
print(f"\n{n_correct}/4 correct")

Run it. Expected output:

epoch     0  loss 0.2501
epoch  1000  loss 0.2500
epoch  2000  loss 0.2500
epoch  3000  loss 0.2500
epoch  4000  loss 0.2500
epoch  5000  loss 0.2500

sensor_a  sensor_b  target  predicted  rounded  result
    0        0       0     0.5000      0      OK
    0        1       1     0.5000      0      WRONG
    1        0       1     0.5000      0      WRONG
    1        1       0     0.5000      0      OK

2/4 correct

Notice the loss does not budge. It starts at 0.2501 and is still 0.2500 five thousand epochs later; the neuron has effectively given up and settled on predicting almost exactly 0.5 (a coin flip) for every input, because that is the one output that minimizes average squared error when the model genuinely cannot find a real pattern to separate the classes. Rounding a near-0.5 guess to a whole number is basically a coin flip too, and in this run it happened to land on 0 every time, which is why exactly half the rows say OK and half say WRONG. This is not a bug. It is precisely what the geometry above predicted: no single straight line separates these four points, so no single neuron ever will, regardless of training time or learning rate.

Step 4: Add a Hidden Layer

The fix is to give the network room to combine more than one straight-line boundary. Add a middle (“hidden”) layer of several neurons between the inputs and the output. Each hidden neuron draws its own straight decision boundary using its own weights and bias, and the output neuron then learns to combine those boundaries into a shape that is no longer a single straight line, which is exactly what solving XOR requires.

One detail matters here: the activation function has to stay non-linear at every layer for this trick to work. If you removed the sigmoid and used the raw weighted sum instead, the algebra collapses. Feeding z1 = X.W1 + b1 straight into a second layer as z2 = z1.W2 + b2 is algebraically identical to z2 = X.(W1.W2) + (b1.W2 + b2), which is just one big linear transformation wearing a two-layer disguise. Stacking any number of layers without a non-linear activation function between them is mathematically the same as a single layer. The non-linearity is what actually buys the network the ability to solve problems like XOR.

Forward propagation through two layers works exactly like Step 2, just twice: compute the hidden layer’s weighted sum and activation first, then feed that hidden layer’s output into the output neuron as if it were the input.

Backpropagation gets one more step too. The output neuron’s error is calculated the same way as before. Then, to find each hidden neuron’s share of the blame, the output layer’s error signal is projected backward through the output layer’s weights and multiplied by the hidden layer’s own activation slope. This is the chain rule again, just applied one layer further back; the same idea as Step 2, only unrolled twice instead of once.

Save this as hidden_layer_network.py:

import numpy as np

np.random.seed(42)


def sigmoid(z):
    return 1 / (1 + np.exp(-z))


X = np.array([[0, 0], [0, 1], [1, 0], [1, 1]], dtype=float)
y = np.array([[0], [1], [1], [0]], dtype=float)

n_inputs, n_hidden, n_outputs = 2, 4, 1

W1 = np.random.randn(n_inputs, n_hidden) * 0.5
b1 = np.zeros((1, n_hidden))
W2 = np.random.randn(n_hidden, n_outputs) * 0.5
b2 = np.zeros((1, n_outputs))

lr = 1.0
epochs = 10000

for epoch in range(epochs):
    # forward propagation
    z1 = X @ W1 + b1
    a1 = sigmoid(z1)
    z2 = a1 @ W2 + b2
    a2 = sigmoid(z2)

    loss = np.mean((a2 - y) ** 2)

    # backpropagation
    d_loss_d_a2 = 2 * (a2 - y) / len(X)
    d_a2_d_z2 = a2 * (1 - a2)
    delta2 = d_loss_d_a2 * d_a2_d_z2

    dW2 = a1.T @ delta2
    db2 = np.sum(delta2, axis=0, keepdims=True)

    d_a1_d_z1 = a1 * (1 - a1)
    delta1 = (delta2 @ W2.T) * d_a1_d_z1

    dW1 = X.T @ delta1
    db1 = np.sum(delta1, axis=0, keepdims=True)

    W2 -= lr * dW2
    b2 -= lr * db2
    W1 -= lr * dW1
    b1 -= lr * db1

    if epoch % 2000 == 0:
        print(f"epoch {epoch:5d}  loss {loss:.4f}")

print(f"epoch {epochs:5d}  loss {loss:.4f}")
print()
print("sensor_a  sensor_b  target  predicted  rounded  result")
z1 = X @ W1 + b1
a1 = sigmoid(z1)
z2 = a1 @ W2 + b2
a2 = sigmoid(z2)
n_correct = 0
for xi, yi, ai in zip(X, y, a2):
    pred = round(ai[0])
    ok = pred == yi[0]
    n_correct += ok
    print(f"    {int(xi[0])}        {int(xi[1])}       {int(yi[0])}     {ai[0]:.4f}      {pred}      {'OK' if ok else 'WRONG'}")
print(f"\n{n_correct}/4 correct")

Run it. Expected output:

epoch     0  loss 0.2557
epoch  2000  loss 0.0052
epoch  4000  loss 0.0010
epoch  6000  loss 0.0005
epoch  8000  loss 0.0004
epoch 10000  loss 0.0003

sensor_a  sensor_b  target  predicted  rounded  result
    0        0       0     0.0187      0      OK
    0        1       1     0.9845      1      OK
    1        0       1     0.9844      1      OK
    1        1       0     0.0151      0      OK

4/4 correct

The loss drops from about 0.26 down to 0.0003 over 10,000 epochs, and all four predictions now round to the correct target. The four hidden neurons each learned their own slice of the decision boundary (print W1 yourself and you will see no two of its four columns match), and the output neuron learned how to combine those four boundaries into the non-linear shape XOR actually needs.

Common Mistake: Initializing All Weights to Zero

It is tempting to initialize weights to zero instead of small random values; zero feels like a neutral, unbiased starting point. Try it and the network gets permanently stuck. Change only the initialization lines from Step 4:

# The mistake: zeros instead of small random values.
W1 = np.zeros((n_inputs, n_hidden))
b1 = np.zeros((1, n_hidden))
W2 = np.zeros((n_hidden, n_outputs))
b2 = np.zeros((1, n_outputs))

Everything else about the script stays identical. Expected output after the same 10,000 epochs:

epoch     0  loss 0.2500
epoch  2000  loss 0.2500
epoch  4000  loss 0.2500
epoch  6000  loss 0.2500
epoch  8000  loss 0.2500
epoch 10000  loss 0.2500

W1 after 10,000 epochs (each column is one hidden neuron):
[[0. 0. 0. 0.]
 [0. 0. 0. 0.]]

sensor_a  sensor_b  target  predicted  rounded  result
    0        0       0     0.5000      0      OK
    0        1       1     0.5000      0      WRONG
    1        0       1     0.5000      0      WRONG
    1        1       0     0.5000      0      OK

2/4 correct

W1 is still exactly all zeros after 10,000 epochs of training, and the network fails identically to the single neuron from Step 3. This is called the symmetry problem. When every weight starts at zero, all four hidden neurons compute the exact same weighted sum, the exact same activation, and therefore the exact same gradient at every single step. They update by identical amounts forever and never differentiate from each other, so a “4-hidden-neuron” network with zero initialization behaves like it only has one useful neuron, which is exactly as powerless against XOR as Step 3’s single neuron was. Small random values break that symmetry by giving each neuron a different starting point, so backpropagation pushes each one in a slightly different direction.

Common Mistake: Picking a Learning Rate Blindly

The learning rate (lr in the code above) controls how big a step gradient descent takes on each update. It is not a value you can guess once and forget; too small and training crawls, too large and it can overshoot the same way. Wrap Step 4’s training loop in a function and sweep it across four learning rates, save this as lr_sweep.py:

import numpy as np


def sigmoid(z):
    return 1 / (1 + np.exp(-np.clip(z, -500, 500)))


def train(lr, epochs=10000, seed=42):
    np.random.seed(seed)
    X = np.array([[0, 0], [0, 1], [1, 0], [1, 1]], dtype=float)
    y = np.array([[0], [1], [1], [0]], dtype=float)
    n_inputs, n_hidden, n_outputs = 2, 4, 1
    W1 = np.random.randn(n_inputs, n_hidden) * 0.5
    b1 = np.zeros((1, n_hidden))
    W2 = np.random.randn(n_hidden, n_outputs) * 0.5
    b2 = np.zeros((1, n_outputs))
    losses = []
    for epoch in range(epochs):
        z1 = X @ W1 + b1
        a1 = sigmoid(z1)
        z2 = a1 @ W2 + b2
        a2 = sigmoid(z2)
        loss = np.mean((a2 - y) ** 2)
        losses.append(loss)
        d_loss_d_a2 = 2 * (a2 - y) / len(X)
        d_a2_d_z2 = a2 * (1 - a2)
        delta2 = d_loss_d_a2 * d_a2_d_z2
        dW2 = a1.T @ delta2
        db2 = np.sum(delta2, axis=0, keepdims=True)
        d_a1_d_z1 = a1 * (1 - a1)
        delta1 = (delta2 @ W2.T) * d_a1_d_z1
        dW1 = X.T @ delta1
        db1 = np.sum(delta1, axis=0, keepdims=True)
        W2 -= lr * dW2
        b2 -= lr * db2
        W1 -= lr * dW1
        b1 -= lr * db1
    return losses


for lr in [0.01, 1.0, 10.0, 50.0]:
    losses = train(lr)
    print(f"lr={lr:>5}: loss at epoch 0={losses[0]:.4f}  epoch 2000={losses[2000]:.4f}  epoch 10000={losses[-1]:.4f}")

Run it. Expected output:

lr= 0.01: loss at epoch 0=0.2557  epoch 2000=0.2503  epoch 10000=0.2501
lr=  1.0: loss at epoch 0=0.2557  epoch 2000=0.0052  epoch 10000=0.0003
lr= 10.0: loss at epoch 0=0.2557  epoch 2000=0.0001  epoch 10000=0.0000
lr= 50.0: loss at epoch 0=0.2557  epoch 2000=0.2500  epoch 10000=0.2499

At lr=0.01, training has barely moved after 10,000 epochs; the loss is still sitting near 0.25, indistinguishable from a genuinely unsolvable problem unless you already know better. At lr=50.0, the steps are so large that training overshoots the good solution on every update and lands back near the same stuck point, again around 0.25. Both failure modes produce the same symptom (loss stuck near 0.25) for opposite reasons, which is exactly why “is my learning rate reasonable” deserves to be one of the first questions you ask when training refuses to converge, right alongside “is this problem actually solvable by this architecture.” In this specific case, lr=1.0 and lr=10.0 both work well; that is not a general rule, it is a property of this tiny 4-input dataset, and real projects tune the learning rate empirically rather than trusting a single fixed number.

Step 5: Compare With a Real Framework (PyTorch)

Everything above was implemented by hand specifically so you would understand what is happening. In practice, almost nobody hand-writes backpropagation for production models; frameworks like PyTorch compute the same gradients automatically (a feature called autograd) and provide pre-built layers and optimizers. Install it if you want to follow along (this adds a meaningfully larger download than NumPy):

pip install torch

This tutorial was tested against torch 2.13.0+cpu. Save this as pytorch_equivalent.py. It builds the identical 2-input, 4-hidden-neuron, 1-output architecture, trained on the identical sensor-disagreement data for the identical 10,000 epochs at the identical learning rate:

import torch
import torch.nn as nn

torch.manual_seed(42)

X = torch.tensor([[0., 0.], [0., 1.], [1., 0.], [1., 1.]])
y = torch.tensor([[0.], [1.], [1.], [0.]])

model = nn.Sequential(
    nn.Linear(2, 4),
    nn.Sigmoid(),
    nn.Linear(4, 1),
    nn.Sigmoid(),
)

loss_fn = nn.MSELoss()
optimizer = torch.optim.SGD(model.parameters(), lr=1.0)

epochs = 10000
for epoch in range(epochs):
    optimizer.zero_grad()
    predictions = model(X)
    loss = loss_fn(predictions, y)
    loss.backward()
    optimizer.step()

    if epoch % 2000 == 0:
        print(f"epoch {epoch:5d}  loss {loss.item():.4f}")

print(f"epoch {epochs:5d}  loss {loss.item():.4f}")
print()
print("sensor_a  sensor_b  target  predicted  rounded  result")
with torch.no_grad():
    predictions = model(X)
n_correct = 0
for xi, yi, pi in zip(X, y, predictions):
    pred = round(pi.item())
    ok = pred == yi.item()
    n_correct += ok
    print(f"    {int(xi[0])}        {int(xi[1])}       {int(yi.item())}     {pi.item():.4f}      {pred}      {'OK' if ok else 'WRONG'}")
print(f"\n{n_correct}/4 correct")
print(f"\nTrainable parameters: {sum(p.numel() for p in model.parameters())}")

Run it. Expected output:

epoch     0  loss 0.2866
epoch  2000  loss 0.0027
epoch  4000  loss 0.0008
epoch  6000  loss 0.0004
epoch  8000  loss 0.0003
epoch 10000  loss 0.0002

sensor_a  sensor_b  target  predicted  rounded  result
    0        0       0     0.0138      0      OK
    0        1       1     0.9833      1      OK
    1        0       1     0.9868      1      OK
    1        1       0     0.0159      0      OK

4/4 correct

Trainable parameters: 17

nn.Linear(2, 4) is PyTorch’s built-in replacement for the W1/b1 pair you wrote by hand: 2 inputs times 4 hidden neurons plus 4 biases is 12 numbers, and nn.Linear(4, 1) similarly replaces W2/b2 with 4 weights plus 1 bias, for 17 trainable numbers total, matching the parameter count of the from-scratch version exactly. PyTorch’s own documentation for nn.MSELoss defines its default behavior (reduction='mean') as the mean of (prediction - target) squared across the batch, the identical formula used by hand above, and loss.backward() is PyTorch’s autograd engine running the same chain rule Step 4 derived manually, just automatically and for arbitrarily deep networks. The final loss (0.0002) and every rounded prediction land in the same place as the hand-written version’s (0.0003), with the small difference coming down to normal floating-point and optimizer-implementation noise, not a different algorithm.

How to Confirm Everything Worked End to End

Before moving on, verify each piece independently:

  • Rerun single_neuron.py (Step 2) and confirm the loss ends near 0.001 with all four OR predictions correct. This proves your sigmoid, forward pass, and gradient descent update are wired correctly on a problem that should be solvable.
  • Rerun single_neuron_fails.py (Step 3) and confirm the loss gets stuck at exactly 0.2500 with only 2 of 4 correct. If it instead solves XOR, double-check you are really using a single neuron with no hidden layer; a single sigmoid neuron mathematically cannot solve XOR, so a passing result there means the architecture is not what you think it is.
  • Rerun hidden_layer_network.py (Step 4) and confirm all 4 predictions are correct with loss near 0.0003. Then print W1 and confirm its four columns are not identical to each other; identical columns mean the symmetry problem from the zero-initialization mistake crept back in.
  • Rerun the PyTorch version (Step 5) and confirm it also reaches 4/4 correct with a comparably small loss and exactly 17 trainable parameters.

Since this tutorial’s dataset is the complete, exhaustive truth table for OR and XOR (there are only 4 possible combinations of two binary inputs, and all 4 were used for training), “4/4 correct” here means the network has fully learned the function, not just memorized a lucky subset. Real projects train on a sample of a much larger space and hold out separate data specifically to check whether the network generalizes to inputs it never saw; this toy problem does not need that step because there is nothing left to hold out.

A Note on the Loss Function, and What to Learn Next

This tutorial used mean squared error (MSE) throughout because it is the simplest loss function to derive and explain by hand. For real classification problems, MSE is not the standard choice; its gradient shrinks as a wrong prediction gets more confidently wrong, which slows learning exactly when you need it most. Binary cross-entropy is the standard loss function for problems like this one, and it is worth learning next specifically because it does not have that weakness.

Other reasonable next steps, roughly in order of how directly they build on what you just did:

  • Swap the MSE loss in this tutorial’s from-scratch code for binary cross-entropy, and derive its gradient by hand the same way Step 2 did for MSE.
  • Train on a real, larger dataset (MNIST handwritten digits is the classic beginner choice) using the same forward propagation and backpropagation pattern you just built, extended to more inputs, more hidden neurons, and more than one output class.
  • Learn mini-batch gradient descent. Every example in this tutorial used full-batch training (all 4 rows, every epoch); real datasets are usually too large for that, and are instead split into small batches per update.
  • Explore other activation functions, particularly ReLU, which has mostly replaced sigmoid in hidden layers of modern networks because it does not suffer from the same vanishing-gradient slowdown on deep networks.
  • Once the fundamentals feel solid, spend time in PyTorch’s own documentation on torch.nn and autograd to see how the concepts from this tutorial (layers, loss functions, gradients, optimizers) map onto a production-grade framework built to run on GPUs and scale to billions of parameters.

Tags:

Machine Learningneural-networksnumpyPythonpytorch

Share

An empty wood-paneled corporate boardroom with a long conference table and rows of chairs
Previous Post

Andreessen Horowitz’s DOJ Probe Turns VC Board Seats Into an Antitrust Test

A giant panda holds bamboo up to its mouth while eating, seated on the ground.
Next Post

ToxicPanda 2.0 Abuses Android VPN Permissions to Blind Google Play Protect

No Comment! Be the first one.

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Latest
08 Oct
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
08 Oct
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
Trending
October 8, 2026
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
October 8, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
October 8, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google

Related Posts

A laptop wrapped in a chain and padlock, illustrating least-privilege controls for AI agents.
Learning Hub

How to Secure Tool-Using AI Agents Before They Touch Production

June 8, 2026
Colorful sticky notes arranged on an office wall, symbolizing governance checklists and planning.
Learning Hub

AI Governance for Agentic Apps: A Practical Checklist for Builders

June 8, 2026
A technician connects green fiber optic cables at a data center, representing a private production inference endpoint.
Learning Hub

How to Deploy a Fine-Tuned LLM Behind a Private Production Inference Endpoint

June 8, 2026
Narrow aisle behind black supercomputer racks in a data center
Learning Hub

Kubernetes SELinux Volume Labeling: What Cluster Operators Should Audit Before v1.37

June 8, 2026
SXZ.io SXZ.io
  • [email protected]

Categories

Articles
Learning Hub
News

All Rights Reserved by SXZ.io ©2026