Have you ever wondered how a computer can recognize a handwritten number, predict whether an email is spam, recommend a video, or understand a sentence?

A lot of modern AI systems rely on something called a neural network.

Now, the name can make them sound much more complicated than they really are. You might imagine that you need advanced calculus, a huge computer, and thousands of lines of code to build one.

You don't.

At its most basic level, a neural network is a mathematical model that takes some numbers as input, performs calculations on those numbers, makes a prediction, checks how far that prediction was from the correct answer, and then adjusts itself so it can do a little better next time.

In this tutorial, we're going to build one ourselves using Python and NumPy.

Here's What We'll Cover:

We'll start with a single artificial neuron, then gradually put together a complete neural network. By the end, you'll understand what weights and biases are, what activation functions do, how a network learns from its mistakes, what backpropagation and gradient descent actually mean, and how all of those pieces fit together.

You don't need to know advanced machine learning to follow along. Some basic Python and algebra will help, but I'll explain the important math as we go.

1. What Is a Neural Network?

Let's start with a simple example.

Imagine that we want a computer to predict whether a student will pass an exam.

We could give the computer information such as:

  • How many hours the student studied

  • How many practice questions they completed

  • Their previous test score

For example:

Study Hours = 5 Practice Questions = 80 Previous Score = 82

We also know whether the student actually passed:

Passed = 1

After seeing many examples like this, we want the computer to learn a pattern.

Maybe students who study more tend to perform better. Maybe previous test scores are useful. Maybe practice questions are helpful, too.

Instead of writing all of those rules ourselves, we can give the examples to a neural network and let it learn the relationships.

The basic idea looks like this:

Visual idea about how a neural network works

The prediction could be something like: 0.92

If we're predicting the probability of passing, we could interpret that as approximately a 92% predicted chance of passing.

The important thing is that we didn't tell the network that...

"Study hours are important, and previous scores are slightly more important."

Instead, the network learns numbers called weights that determine how strongly different inputs affect its predictions.

2. Why Are They Called Neural Networks?

The name comes from biological brains.

Your brain contains neurons that receive signals, process information, and pass signals to other neurons.

Artificial neural networks are not artificial brains. They don't work exactly like biological neurons. But the general idea of connecting many simple processing units inspired the name.

A very simplified artificial neuron looks like this:

Input and Output through a neural network

The neuron receives numbers, performs some mathematical operations, and produces another number.

A neural network is made by connecting many of these artificial neurons together.

3. The Three Main Parts of a Neural Network

A simple neural network can be divided into three types of layers:

  1. Input Layer

  2. Hidden Layer(s)

  3. Output Layer

Let's look at each one.

The Input Layer

The input layer contains the information we give the network.

For our student example, we could have three inputs:

Input 1 = Study Hours
Input 2 = Practice Questions
Input 3 = Previous Score

So one student's input might look like:

[5, 80, 82]

The network doesn't necessarily understand that these numbers mean "study hours" or "test score." To the mathematical part of the network, they're simply numbers.

That's an important idea to remember:

Neural networks work with numbers.

Images, text, audio, and other information must eventually be represented as numbers before a neural network can process them.

Hidden Layers

After the input layer come the hidden layers.

A network might look like:

Image showing how data moves from the Input Layer to the Hidden Layer and then to the Output layer

The hidden layer contains neurons that perform calculations on the inputs.

A network can have one hidden layer or many hidden layers.

When a network has many layers, we often call it a deep neural network.

The Output Layer

The output layer produces the final result.

For a simple yes/no problem, we might represent the answers as:

0 = No
1 = Yes

For example:

0.12 → probably No
0.91 → probably Yes

For a problem with multiple categories, the output could contain several numbers:

Cat  = 0.05
Dog  = 0.90
Bird = 0.05

The largest value is associated with "Dog," so the model would predict Dog.

4. What Is a Neuron?

Now let's zoom in on one neuron.

Suppose our neuron receives three inputs:

x₁
x₂
x₃

Each input has a corresponding weight:

w₁
w₂
w₃

The neuron multiplies each input by its weight and adds the results together.

It also adds something called a bias.

The equation is:

z = x₁w₁ + x₂w₂ + x₃w₃ + b

Don't worry if that equation looks intimidating.

It's basically just:

input × weight
+
input × weight
+
input × weight
+
bias

Let's use actual numbers.

Suppose:

x₁ = 2
x₂ = 3
x₃ = 4

w₁ = 0.5
w₂ = 0.2
w₃ = 0.8

b = 1

Then:

z = (2 × 0.5) + (3 × 0.2) + (4 × 0.8) + 1

Calculate each part:

2 × 0.5 = 1.0
3 × 0.2 = 0.6
4 × 0.8 = 3.2

Now add them:

z = 1.0 + 0.6 + 3.2 + 1
z = 5.8

The neuron has produced 5.8.

But we're not finished yet.

5. What Is a Weight?

A weight controls how strongly an input affects a neuron.

Imagine we have:

x = 5

If the weight is:

w = 2

then:

x × w = 5 × 2
      = 10

But if the weight is:

w = 0.1

then:

x × w = 5 × 0.1
      = 0.5

The same input produced a very different result because the weight changed.

You can think of a weight as a volume knob.

A large positive weight makes an input have a stronger positive influence. A weight close to zero makes the input have little influence. A negative weight can push the result in the opposite direction.

The network learns these weights during training.

6. What Is a Bias?

The bias is another number added to the neuron's calculation.

Without the bias, we would have:

z = x₁w₁ + x₂w₂ + x₃w₃

With the bias:

z = x₁w₁ + x₂w₂ + x₃w₃ + b

Why add another number? Because it gives the neuron more flexibility.

Think of it like adjusting the starting point of the neuron's calculation.

The network learns the bias during training just like it learns the weights.

So when you see:

weights + bias

you're looking at some of the parameters the neural network can change while it learns.

7. Why Do We Need Activation Functions?

At this point, our neuron can calculate a weighted sum:

z = x₁w₁ + x₂w₂ + ... + b

But neural networks need to learn more complicated relationships than simple weighted sums.

That's where activation functions come in. An activation function takes the neuron's calculated value and transforms it.

One common activation function is ReLU. ReLU stands for Rectified Linear Unit.

Its equation is:

ReLU(x) = max(0, x)

In simple terms:

  • If the number is positive, keep it.

  • If the number is negative, turn it into zero.

For example:

ReLU(-5) = 0
ReLU(-2) = 0
ReLU(0)  = 0
ReLU(3)  = 3
ReLU(10) = 10

In Python:

def relu(x):
    return max(0, x)

With NumPy arrays, we can use:

def relu(x):
    return np.maximum(0, x)

Activation functions are important because they allow neural networks with multiple layers to learn more complicated patterns.

8. Building Our First Neuron in Python

Let's turn the math into Python.

First, import NumPy:

import numpy as np

NumPy gives us tools for working with numbers, arrays, vectors, and matrices.

Now let's create our inputs:

x = np.array([2, 3, 4])

This creates an array containing three values:

[2, 3, 4]

Now create the weights:

weights = np.array([0.5, 0.2, 0.8])

We have one weight for each input:

x₁ = 2    w₁ = 0.5
x₂ = 3    w₂ = 0.2
x₃ = 4    w₃ = 0.8

Next, create the bias:

bias = 1

Now we calculate the weighted sum:

z = np.dot(x, weights) + bias

np.dot() performs the multiplication-and-addition operation we described earlier.

In this case:

np.dot(x, weights)

is equivalent to:

(2 × 0.5) + (3 × 0.2) + (4 × 0.8)

which equals:

4.8

Then we add the bias:

4.8 + 1 = 5.8

Now apply ReLU:

output = np.maximum(0, z)

Since z is 5.8, ReLU leaves it unchanged:

output = 5.8

Finally:

print(output)

prints:

5.8

So our entire neuron is:

import numpy as np

x = np.array([2, 3, 4])
weights = np.array([0.5, 0.2, 0.8])
bias = 1

z = np.dot(x, weights) + bias
output = np.maximum(0, z)

print(output)

We have just created a tiny artificial neuron.

9. From One Neuron to a Layer

One neuron isn't enough for most interesting problems.

Instead, we can connect several neurons together.

For example:

Input, Output and Hidden Layer depicted with neurons

Those neurons together form a layer.

A small neural network might look like:

Input Layer
     ↓
Hidden Layer
     ↓
Output Layer

Every neuron in one layer can send its output to neurons in the next layer.

This is where neural networks start becoming much more powerful.

10. How Does a Neural Network Actually Learn?

So far, we've manually chosen the weights:

0.5
0.2
0.8

But a real neural network doesn't start out knowing the correct weights.

Instead, it starts with weights that are usually initialized to small random values.

Then it goes through a cycle:

Make a prediction
       ↓
Compare prediction with correct answer
       ↓
Measure the error
       ↓
Figure out how to change the weights
       ↓
Update the weights
       ↓
Try again

This process happens over and over, and the network gradually adjusts its parameters to make better predictions on the training data.

Let's break each part down.

11. Predictions and Loss

Suppose the correct answer is:

1

but our network predicts:

0.3

The prediction isn't very close to the target.

We need a way to measure how wrong it is. That's what a loss function does.

A loss function takes the prediction and the correct answer and produces a number representing the model's error.

For a simple example, we could use squared error:

Loss = (prediction - actual)²

Using our numbers:

Loss = (0.3 - 1)²

First:

0.3 - 1 = -0.7

Then square it:

(-0.7)² = 0.49

So:

Loss = 0.49

Generally, a smaller loss means the prediction is closer to the target.

In real neural networks, different problems use different loss functions. For binary classification, binary cross-entropy is commonly used.

12. What Are Gradients?

Now we have a problem.

We know that the prediction was wrong, but how should we change the weights?

This is where gradients become useful. A gradient tells us how changing a parameter would affect the loss.

You can think of it like standing on a hill. Imagine that your goal is to reach the lowest point. If you know which direction slopes upward, you can move in the opposite direction to go downhill.

Training a neural network works with a similar idea. We want to reduce the loss. The gradients give us information about which direction the parameters should move.

13. What Is Gradient Descent?

Gradient descent is the process of using gradients to adjust the network's parameters.

A simplified update rule is:

new weight = old weight - learning rate × gradient

In Python:

weight = weight - learning_rate * gradient

The learning rate controls how large the update is.

For example:

learning_rate = 0.01

If the learning rate is too large, the network can make huge changes and potentially jump around instead of settling on a good solution.

If it's too small, learning can take a very long time.

So training involves finding parameter updates that move the model toward lower loss without making the process unstable.

14. What Is Backpropagation?

There's still one important question:

If a neural network has thousands or millions of weights, how does it figure out which weights contributed to the error?

That's where backpropagation comes in. Backpropagation calculates gradients for the parameters by working backward through the network.

Imagine a network like this:

Input
  ↓
Hidden Layer
  ↓
Output
  ↓
Loss

During the forward pass, information moves:

Input → Hidden Layer → Output

During backpropagation, gradient information moves backward:

Loss → Output → Hidden Layer → Input

The network uses these gradients to determine how its weights and biases should change.

You don't normally calculate all of these derivatives by hand when building real neural networks. Libraries such as PyTorch can calculate them automatically.

But understanding the basic idea is important:

Backpropagation calculates how the parameters contributed to the error, and gradient descent uses that information to update them.

15. The Complete Learning Cycle

Now we can put everything together.

Learning cycle of neural network: input, prediction, loss, gradients, update (and then back to prediction...)

More specifically:

Give the network data
          ↓
Calculate a prediction
          ↓
Compare it with the correct answer
          ↓
Calculate the loss
          ↓
Calculate gradients
          ↓
Update weights and biases
          ↓
Repeat

One complete pass through the training data is often called an epoch.

For example:

Epoch 1 → Loss: 0.82
Epoch 2 → Loss: 0.61
Epoch 3 → Loss: 0.43
Epoch 4 → Loss: 0.29
Epoch 5 → Loss: 0.18

These numbers are just an example, but ideally the loss decreases as training progresses.

16. Let's Build a Neural Network From Scratch

Congrats! You now understand the basics of neural networks. Now it's time to put these ideas together.

We're going to build a small neural network using only:

Python + NumPy

Our network will learn a classic machine learning problem called XOR.

XOR is a logical operation with two inputs.

Its rules are:

0 XOR 0 → 0
0 XOR 1 → 1
1 XOR 0 → 1
1 XOR 1 → 0

In other words, the output is 1 when exactly one of the inputs is 1.

Our training data will therefore be:

X = np.array([
    [0, 0],
    [0, 1],
    [1, 0],
    [1, 1]
])

And the correct answers are:

y = np.array([
    [0],
    [1],
    [1],
    [0]
])

We want our neural network to learn this pattern.

17. Understanding the Network Architecture

Our network will contain:

2 input neurons
       ↓
4 hidden neurons
       ↓
1 output neuron

The two inputs represent the two numbers in each XOR example.

The four hidden neurons give the network enough flexibility to learn the XOR relationship.

The output neuron produces a number between 0 and 1.

18. Setting Up the Data

Let's start our Python program.

import numpy as np

This imports NumPy. We'll use NumPy for arrays, matrix multiplication, and mathematical operations.

Next:

X = np.array([
    [0, 0],
    [0, 1],
    [1, 0],
    [1, 1]
])

X contains our four training examples.

Each row is one example:

[0, 0]
[0, 1]
[1, 0]
[1, 1]

Now create the correct answers:

y = np.array([
    [0],
    [1],
    [1],
    [0]
])

The first row of X corresponds to the first row of y.

So:

[0, 0] → 0
[0, 1] → 1
[1, 0] → 1
[1, 1] → 0

19. Creating the Weights and Biases

Now we need the parameters of our network.

First:

np.random.seed(42)

This makes our random numbers reproducible.

Without this line, the network would receive different random starting weights each time we ran the program.

Now create the first layer's weights:

W1 = np.random.randn(2, 4)

Why (2, 4)?

Because:

  • We have 2 input values.

  • We have 4 neurons in the hidden layer.

So W1 needs a weight connecting each input to each hidden neuron.

There are:

2 × 4 = 8

weights.

Next:

b1 = np.zeros((1, 4))

This creates four biases, one for each hidden neuron.

Now the second layer:

W2 = np.random.randn(4, 1)

There are four hidden neurons and one output neuron, so we need:

4 × 1 = 4

weights.

Finally:

b2 = np.zeros((1, 1))

This gives the output neuron one bias.

Our network parameters are therefore:

W1 → input-to-hidden weights
b1 → hidden-layer biases

W2 → hidden-to-output weights
b2 → output-layer bias

20. The Sigmoid Function

Our output represents a probability, so we'd like it to be between 0 and 1.

We can use the sigmoid function.

Its equation is:

sigmoid(x) = 1 / (1 + e⁻ˣ)

In Python:

def sigmoid(x):
    return 1 / (1 + np.exp(-x))

Let's see what it does:

sigmoid(-5) ≈ 0.007
sigmoid(0)  = 0.5
sigmoid(5)  ≈ 0.993

No matter how large or small the input is, the result stays between 0 and 1.

That's useful when our output represents a probability.

21. Forward Propagation

Now we can send the data through the network. This is called forward propagation.

First, calculate the hidden layer:

z1 = X @ W1 + b1

There's a new symbol here:

@

In Python, @ performs matrix multiplication.

You can think of this operation as performing many weighted sums at once.

Instead of manually calculating every neuron:

input × weight + input × weight + bias

NumPy can calculate all of them together.

The result is stored in z1.

Next:

a1 = np.tanh(z1)

Here we're using the tanh activation function for the hidden layer.

Tanh converts its input into values between -1 and 1.

Why use tanh here?

Because XOR isn't something a single simple linear calculation can solve. The nonlinear activation gives the hidden layer the flexibility it needs to learn the pattern.

Now calculate the output layer:

z2 = a1 @ W2 + b2

This takes the hidden layer's outputs and combines them using the second set of weights.

Finally:

a2 = sigmoid(z2)

Now a2 contains our predictions.

For example, before training, the network might produce something like:

0.52
0.61
0.48
0.55

Those predictions aren't useful yet, but that's expected. The network hasn't learned anything yet.

22. Calculating the Loss

Now we need to measure how good those predictions are.

For binary classification, we'll use binary cross-entropy, which is a loss function used in machine learning for binary classification. It measures the performance of a model whose output is a probability value between 0 and 1.

The formula is:

Loss = -mean(
    y × log(prediction)
    +
    (1 - y) × log(1 - prediction)
)

That looks much more complicated than the squared-error example from earlier, but we don't need to memorize the formula.

In Python:

loss = -np.mean(
    y * np.log(a2 + 1e-8) +
    (1 - y) * np.log(1 - a2 + 1e-8)
)

The 1e-8 is a very small number.

It prevents problems if a2 gets extremely close to 0 or 1, because taking the logarithm of exactly zero isn't valid.

At the beginning of training, the loss will probably be relatively high. But as the network learns, we'd like it to decrease.

23. Backpropagation in Code

Now comes the most mathematical part of our program.

We need to calculate the gradients.

Start with:

dz2 = a2 - y

This gives us the gradient of the loss with respect to the output layer's pre-activation value for the sigmoid + binary cross-entropy combination.

Next:

dW2 = (a1.T @ dz2) / len(X)

This calculates the gradient for W2.

The .T means transpose.

Our hidden-layer output has four neurons, while dz2 represents the output layer's error. Matrix multiplication combines them to determine how each hidden-to-output weight contributed to the loss.

We divide by:

len(X)

because we have four training examples and we're calculating the average gradient.

Now calculate the output bias gradient:

db2 = np.mean(dz2, axis=0, keepdims=True)

This calculates the average gradient for the output bias.

Next:

da1 = dz2 @ W2.T

This sends the gradient information backward from the output layer toward the hidden layer.

Now we need to account for the derivative of the tanh activation function.

The derivative of tanh can be written as:

1 - tanh(x)²

Since we already have the hidden layer's activated values in a1, we can write:

dz1 = da1 * (1 - a1**2)

This tells us how the hidden layer's pre-activation values affected the loss.

Now calculate the gradients for the first layer's weights:

dW1 = (X.T @ dz1) / len(X)

And the hidden-layer biases:

db1 = np.mean(dz1, axis=0, keepdims=True)

At this point, we have gradients for all of our trainable parameters.

24. Updating the Weights

Now we use gradient descent.

First:

W2 -= learning_rate * dW2

This updates the second layer's weights.

The -= means:

W2 = W2 - learning_rate * dW2

Then:

b2 -= learning_rate * db2

updates the output bias.

And:

W1 -= learning_rate * dW1

updates the first layer's weights.

Finally:

b1 -= learning_rate * db1

updates the hidden-layer biases.

These updates are what actually allow the network to learn.

25. The Complete NumPy Neural Network

Now let's put everything together.

import numpy as np

# 1. Training data

X = np.array([
    [0, 0],
    [0, 1],
    [1, 0],
    [1, 1]
])

y = np.array([
    [0],
    [1],
    [1],
    [0]
])

# 2. Initialize parameters

np.random.seed(42)

W1 = np.random.randn(2, 4)
b1 = np.zeros((1, 4))

W2 = np.random.randn(4, 1)
b2 = np.zeros((1, 1))

learning_rate = 0.1

# 3. Activation functions

def sigmoid(x):
    return 1 / (1 + np.exp(-x))

# 4. Training

for epoch in range(10000):

    # Forward propagation

    z1 = X @ W1 + b1
    a1 = np.tanh(z1)

    z2 = a1 @ W2 + b2
    a2 = sigmoid(z2)

    # Calculate loss

    loss = -np.mean(
        y * np.log(a2 + 1e-8) +
        (1 - y) * np.log(1 - a2 + 1e-8)
    )

    # Backpropagation

    dz2 = a2 - y

    dW2 = (a1.T @ dz2) / len(X)
    db2 = np.mean(dz2, axis=0, keepdims=True)

    da1 = dz2 @ W2.T

    dz1 = da1 * (1 - a1**2)

    dW1 = (X.T @ dz1) / len(X)
    db1 = np.mean(dz1, axis=0, keepdims=True)

    # Update parameters

    W2 -= learning_rate * dW2
    b2 -= learning_rate * db2

    W1 -= learning_rate * dW1
    b1 -= learning_rate * db1

    # Display progress

    if epoch % 1000 == 0:
        print(f"Epoch {epoch}, Loss: {loss:.4f}")

Let's go through the program from top to bottom.

Line-by-Line Explanation of the Full Code

Importing NumPy:

import numpy as np

We import NumPy because our network will work with arrays and matrix operations.

Creating the inputs

X = np.array([
    [0, 0],
    [0, 1],
    [1, 0],
    [1, 1]
])

Each row is one XOR example.

There are four examples and two input values per example.

So the shape of X is:

4 × 2

Creating the answers

y = np.array([
    [0],
    [1],
    [1],
    [0]
])

There are four correct answers, one for each row in X.

Making random initialization reproducible

np.random.seed(42)

This makes NumPy generate the same starting random values each time.

The number 42 isn't special. You could use another number.

Creating the first weight matrix

W1 = np.random.randn(2, 4)

This creates a matrix containing random numbers.

Its shape is 2*4

There are two inputs and four hidden neurons.

Creating the first biases

b1 = np.zeros((1, 4))

This creates four zeros:

[0, 0, 0, 0]

There is one bias for every hidden neuron.

Creating the second weight matrix

W2 = np.random.randn(4, 1)

There are four hidden neurons and one output neuron.

Therefore:

4 × 1

weights are needed.

Creating the output bias

b2 = np.zeros((1, 1))

The output layer has one neuron, so it needs one bias.

Setting the learning rate

learning_rate = 0.1

This controls how strongly the gradients affect each update.

Creating sigmoid

def sigmoid(x):
    return 1 / (1 + np.exp(-x))

This converts the output into a value between 0 and 1.

Starting the Training Loop

for epoch in range(10000):

This tells Python to repeat the training process 10,000 times.

Each repetition is an epoch, which is one complete pass of the entire training dataset through a neural network

Calculating the hidden layer

z1 = X @ W1 + b1

This performs the weighted-sum calculation for all four hidden neurons and all four training examples.

Applying tanh

a1 = np.tanh(z1)

This applies the nonlinear activation function to the hidden layer.

Calculating the output layer

z2 = a1 @ W2 + b2

This takes the hidden layer's values and calculates the output neuron's weighted sum.

Applying sigmoid

a2 = sigmoid(z2)

This turns the output into probabilities between 0 and 1.

Calculating the loss

loss = -np.mean(
    y * np.log(a2 + 1e-8) +
    (1 - y) * np.log(1 - a2 + 1e-8)
)

This measures how different the predictions are from the correct answers.

A lower value generally means the predictions are better.

Calculating the output gradient

dz2 = a2 - y

This calculates the gradient needed to update the output layer.

Updating the second-layer weight gradients

dW2 = (a1.T @ dz2) / len(X)

This determines how each weight connecting the hidden layer to the output layer contributed to the loss.

Updating the output bias gradient

db2 = np.mean(dz2, axis=0, keepdims=True)

This calculates the average gradient for the output bias.

Moving backward toward the hidden layer

da1 = dz2 @ W2.T

This passes the gradient information backward through the output layer.

Applying the tanh derivative

dz1 = da1 * (1 - a1**2)

This accounts for the effect of the tanh activation function.

Calculating the first-layer gradients

dW1 = (X.T @ dz1) / len(X)

This determines how the input-to-hidden weights contributed to the loss.

Then:

db1 = np.mean(dz1, axis=0, keepdims=True)

calculates the gradients for the hidden-layer biases.

Updating the parameters

W2 -= learning_rate * dW2
b2 -= learning_rate * db2

W1 -= learning_rate * dW1
b1 -= learning_rate * db1

These four lines are where the network changes what it has learned.

The gradients tell us which direction to move, while the learning rate determines how large the movement should be.

Printing the loss

if epoch % 1000 == 0:
    print(f"Epoch {epoch}, Loss: {loss:.4f}")

The % operator gives us the remainder after division.

So:

epoch % 1000 == 0

is true every 1,000 epochs.

That means we don't print something 10,000 times. Instead, we get occasional updates such as:

Epoch 0, Loss: ...
Epoch 1000, Loss: ...
Epoch 2000, Loss: ...
...

If training is working well, the loss should generally decrease.

26. Testing the Network

After training, we can use the network to make predictions.

z1 = X @ W1 + b1
a1 = np.tanh(z1)

z2 = a1 @ W2 + b2
predictions = sigmoid(z2)

print(predictions)

The network should produce values close to:

[[0],
 [1],
 [1],
 [0]]

The actual values probably won't be exactly 0 and 1.

You might get something more like:

[[0.01],
 [0.98],
 [0.99],
 [0.02]]

That's fine.

The network is producing probabilities.

We can convert those probabilities into classes using a threshold:

classes = (predictions >= 0.5).astype(int)

print(classes)

The result should be:

[[0],
 [1],
 [1],
 [0]]

Our network has learned the XOR pattern.

27. Why Did We Need a Hidden Layer?

You might wonder why we couldn't just connect the two inputs directly to the output.

The reason is that XOR isn't something a single linear layer can represent.

The hidden layer gives the network additional transformations that allow it to learn the more complicated relationship.

This is one of the most important ideas behind neural networks: a network doesn't necessarily learn one giant rule. Instead, different layers can transform information step by step.

For an image recognition system, you can imagine a simplified process like:

Image recognition system visually depicted

Real neural networks don't literally create neat layers called "edges," "shapes," and "objects." This is just an intuition for how increasingly complex representations can emerge through multiple layers.

28. What Happens in a Larger Neural Network?

The network we built is tiny. Modern neural networks can have millions, billions, or even more parameters.

A simplified network might look like:

Simplified neural network visually depicted

Each connection can have its own weight.

The more neurons and connections a network has, the more parameters it may need to learn.

Large models therefore require significant amounts of computing power and memory.

But remember the basic process:

Input
 ↓
Calculations
 ↓
Prediction
 ↓
Loss
 ↓
Gradients
 ↓
Parameter Updates

The size of the network changes dramatically, but the basic training idea remains.

29. Do You Have to Build Neural Networks From Scratch?

No. Building a neural network from scratch is useful for learning because it forces you to understand what's happening underneath the libraries.

But you normally wouldn't manually calculate every gradient when building a real machine learning application.

That's where machine learning frameworks come in. Some commonly used Python libraries include:

  • NumPy

  • PyTorch

  • TensorFlow

  • Keras

  • scikit-learn

For deep learning, PyTorch is one of the most commonly used frameworks. It can automatically calculate gradients and handle many of the mathematical operations involved in training.

30. Building the Same Network With PyTorch

Let's see how much shorter the network becomes with PyTorch.

First, install it:

pip install torch

Then import it:

import torch
import torch.nn as nn

Now create the model:

model = nn.Sequential(
    nn.Linear(2, 4),
    nn.Tanh(),
    nn.Linear(4, 1),
    nn.Sigmoid()
)

Let's break that down.

nn.Linear(2, 4)

creates a layer that takes two inputs and produces four outputs.

That's our hidden layer.

Next:

nn.Tanh()

applies the tanh activation function.

Then:

nn.Linear(4, 1)

connects the four hidden neurons to one output neuron.

Finally:

nn.Sigmoid()

converts the output into a value between 0 and 1.

So the architecture is:

2 inputs
   ↓
4 hidden neurons
   ↓
Tanh
   ↓
1 output neuron
   ↓
Sigmoid

Notice how much shorter this is than our NumPy implementation.

That's because PyTorch handles many of the calculations for us.

31. Training the Network With PyTorch

First, create the training data:

X = torch.tensor([
    [0., 0.],
    [0., 1.],
    [1., 0.],
    [1., 1.]
])

y = torch.tensor([
    [0.],
    [1.],
    [1.],
    [0.]
])

The decimal points are important because neural networks normally work with floating-point numbers.

Now create the model:

model = nn.Sequential(
    nn.Linear(2, 4),
    nn.Tanh(),
    nn.Linear(4, 1),
    nn.Sigmoid()
)

Next, choose our loss function:

loss_function = nn.BCELoss()

BCELoss calculates binary cross-entropy loss.

Now create an optimizer:

optimizer = torch.optim.Adam(
    model.parameters(),
    lr=0.01
)

Adam is an optimization algorithm that updates the model's parameters during training.

model.parameters() tells the optimizer which values it should update.

lr=0.01 sets the learning rate.

Now we can train:

for epoch in range(5000):

    predictions = model(X)

    loss = loss_function(predictions, y)

    optimizer.zero_grad()

    loss.backward()

    optimizer.step()

    if epoch % 500 == 0:
        print(
            f"Epoch {epoch}, Loss: {loss.item():.4f}"
        )

Let's look at the important parts.

First:

predictions = model(X)

This sends the training data through the network.

Then:

loss = loss_function(predictions, y)

compares the predictions with the correct answers.

Next:

optimizer.zero_grad()

clears gradients from the previous training step.

Then:

loss.backward()

calculates the gradients automatically using backpropagation.

Finally:

optimizer.step()

uses those gradients to update the model's parameters.

That's the same basic learning process we implemented manually with NumPy. The difference is that PyTorch takes care of many of the calculations.

32. NumPy vs. PyTorch

So why did we build the network twice? Well, because the two versions teach different things.

With NumPy, we manually handled weights, biases, forward propagation,
loss, gradients, backpropagation, and parameter updates. That makes the mechanics easier to see.

With PyTorch, we can write the same general idea in much less code because the framework handles many of those calculations.

You can think of it like this:

Comparison between NumPy and PyTorch

Learning how the NumPy version works makes the PyTorch version much less mysterious.

33. What Is Deep Learning?

You may have heard the term deep learning. Deep learning is a part of machine learning that uses neural networks with multiple layers.

For example:

Input
  ↓
Layer 1
  ↓
Layer 2
  ↓
Layer 3
  ↓
Layer 4
  ↓
Output

The word "deep" refers to the depth of the network, or the number of layers involved.

There isn't a magical point where a neural network suddenly becomes intelligent. Adding layers simply gives the model more opportunities to transform the input into useful representations.

34. Where Are Neural Networks Used?

Neural networks are used in many different areas. Here are a few examples...

Computer Vision

Neural networks can process images.

For example:

Image
  ↓
Neural Network
  ↓
Prediction

They can be used for tasks such as image classification and object detection.

Natural Language Processing

Neural networks can also process text.

For example:

Text
  ↓
Neural Network
  ↓
Prediction

Modern language models use neural networks to process and generate text.

Speech Recognition

Neural networks can process audio and help convert spoken language into text.

Audio
  ↓
Neural Network
  ↓
Words

Recommendation Systems

Neural networks can learn patterns from user behavior and help predict which content or products might be useful to someone.

Generative AI

Large neural networks can also be used to generate text, images, audio, code, video, and much more.

These systems are much more complicated than the small XOR network we built, but they still rely on the same general idea of learning parameters from data.

35. The Whole Process in One Picture

At this point, we've covered a lot.

Here's the entire training process:

Data
   ↓
Neural Network
   ↓
Prediction
   ↓
Loss
  ↓
Backpropagation
  ↓
Update Parameters
  ↓
Repeat

Once training is finished, we use the learned parameters to make predictions on new data:

New Data
   ↓
Trained Neural Network
   ↓
Prediction

That's the basic idea behind neural network training.

36. The Most Important Ideas to Remember

If you don't remember every equation from this tutorial, that's okay.

Start with these concepts.

Inputs

The numbers we give to the network.

x₁, x₂, x₃...

Weights

Numbers that determine how strongly inputs affect neurons.

w₁, w₂, w₃...

Biases

Additional values that give neurons more flexibility.

b

Activation Functions

Functions that transform neuron outputs and allow networks to learn nonlinear patterns.

Examples include:

ReLU
Tanh
Sigmoid

Forward Propagation

Sending data from the input toward the output.

Input → Hidden Layers → Output

Loss

A measurement of how different the prediction is from the correct answer.

Backpropagation

Calculating gradients by working backward through the network.

Gradient Descent

Using those gradients to update the network's parameters.

And the entire learning process can be summarized as:

Predict
   ↓
Measure Error
   ↓
Calculate Gradients
   ↓
Update Parameters
   ↓
Repeat

37. What Should You Learn Next?

If you want to continue learning neural networks with Python, you don't need to jump directly into complicated research papers.

A useful learning path is:

Python
  ↓
NumPy
  ↓
Basic Linear Algebra
  ↓
Probability & Statistics
  ↓
Machine Learning Basics
  ↓
Neural Networks
  ↓
PyTorch
  ↓
Deep Learning
  ↓
Computer Vision / NLP / Generative AI

You can also learn by building small projects.

For example:

  1. XOR classifier

  2. House price predictor

  3. Handwritten digit classifier

  4. Simple image classifier

  5. Spam message classifier

  6. Neural network that learns a mathematical function

The projects don't need to be huge. A small project that you completely understand is often more useful than a large project where you copied code without understanding it.

Final Takeaway

Neural networks can look intimidating because the systems used in modern AI can contain enormous numbers of parameters.

But the basic idea is much smaller.

A neural network takes numbers as input, combines them using weights and biases, applies mathematical functions, produces a prediction, measures how wrong that prediction was, and then adjusts its parameters.

The cycle looks like this:

Input
  ↓
Weighted Calculations
  ↓
Activation Functions
  ↓
Prediction
  ↓
Loss
  ↓
Gradients
  ↓
Parameter Updates
  ↓
Repeat

That's the foundation.

The XOR network we built in this tutorial is tiny compared with the neural networks used in modern AI. But the ideas you just learned (parameters, layers, activation functions, forward propagation, loss, backpropagation, gradients, and optimization) are fundamental ideas that appear again and again in deep learning.

The next time you hear that an AI model has millions or billions of parameters, it might still sound overwhelming.

But underneath all that scale, the basic learning loop is still familiar:

  1. Make a prediction.

  2. Measure the error.

  3. Figure out how to improve.

  4. Update the parameters.

  5. Try again.

And that's the core idea behind a neural network.

Happy coding and keep learning!