top of page
Gradient With Circle
Image by Nick Morrison

Insights Across Technology, Software, and AI

Discover articles across technology, software, and AI. From core concepts to modern tech and practical implementations.

Why Is Your PyTorch Model Not Converging? Common Gradient Descent Pitfalls

13 minutes ago
8 min read

Training a neural network often looks deceptively simple: define a model, choose a loss function, create an optimizer, and run the training loop. Yet one of the most frustrating situations in deep learning is watching the loss refuse to decrease. Sometimes it barely changes. Sometimes it oscillates wildly. In other cases, it decreases for a while and then suddenly becomes NaN.


These problems are often not caused by the model architecture itself. Many come from how gradients are calculated and updated during training. Understanding the mechanics of PyTorch Gradient Descent makes these issues much easier to diagnose.


Gradient descent is the process that adjusts model parameters in the direction that reduces the loss. For a parameter θ, the basic update can be represented as:


θ ← θ − η ∇ θL


Here, η is the learning rate and ∇θL is the gradient of the loss with respect to the parameter. PyTorch handles the gradient calculation through its automatic differentiation system, while optimizers such as SGD and Adam perform the parameter updates.


When convergence fails, something in this process may be poorly configured. The learning rate might be inappropriate, gradients might not be reaching the parameters, gradients might be accumulating between iterations, or the loss and model output may not be compatible.


PyTorch Gradient Descent pitfalls

PyTorch Gradient Descent and the Training Loop

Before debugging convergence, it helps to understand what actually happens during one training iteration. A typical PyTorch training loop contains four important operations: clearing old gradients, performing a forward pass, calculating the loss and gradients, and updating the parameters.

The optimizer does not automatically know what the current gradient should be. During the forward pass, PyTorch builds a computational graph. Calling loss.backward() then traverses that graph and calculates derivatives for tensors that require gradients. Finally, optimizer.step() uses those gradients to modify the model parameters. A minimal example looks like this:

import torch
import torch.nn as nn

model = nn.Linear(1, 1)
optimizer = torch.optim.SGD(model.parameters(), lr=0.01)
criterion = nn.MSELoss()

x = torch.tensor([[1.0], [2.0], [3.0]])
y = torch.tensor([[2.0], [4.0], [6.0]])

for epoch in range(100):
    optimizer.zero_grad()

    predictions = model(x)
    loss = criterion(predictions, y)

    loss.backward()
    optimizer.step()

    if epoch % 10 == 0:
        print(f"Epoch {epoch}, Loss: {loss.item():.4f}")

The optimizer.zero_grad() call is particularly important. PyTorch accumulates gradients by default. That behavior is useful in some situations, such as gradient accumulation for large batches, but it is usually not what you want during a standard training iteration. Forgetting to clear the gradients can cause the gradients from previous iterations to be added to the current ones.


For example, if the gradient from one iteration is g₁ and the next iteration produces g₂, the stored gradient can become:


∇L = g₁ + g₂


instead of simply g₂.


As training continues, this can produce increasingly large parameter updates and make the optimization process unstable.

Another common mistake is calling optimizer.step() before loss.backward(). The optimizer needs gradients that have already been calculated. If the gradients are missing or stale, the parameter update cannot represent the current loss.

The order should normally remain:

optimizer.zero_grad()
        ↓
forward pass
        ↓
calculate loss
        ↓
loss.backward()
        ↓
optimizer.step()

The distinction between model.train() and model.eval() can also matter. Layers such as dropout and batch normalization behave differently during training and evaluation. Accidentally leaving a model in evaluation mode during training can therefore produce unexpected results.


Common PyTorch Gradient Descent Pitfalls That Stop Convergence

One of the first things to investigate is the learning rate. The learning rate determines the size of each parameter update. If it is too small, the model may technically be learning but make such tiny updates that training appears to be stuck. If it is too large, the optimizer can repeatedly overshoot useful parameter values. Consider a loss curve that looks like this conceptually:


oscillation in gradient descent

Large oscillations can be a sign that the learning rate is too high. On the other hand, a nearly flat curve can indicate that the learning rate is too low, although vanishing gradients and other problems can produce similar behavior. You can test different learning rates with a simple experiment:

learning_rates = [1e-2, 1e-3, 1e-4]

for lr in learning_rates:
    model = MyModel()
    optimizer = torch.optim.Adam(model.parameters(), lr=lr)

There is no universally correct learning rate. Its useful range depends on the architecture, optimizer, data, initialization, batch size, and loss function. Adam often works reasonably well with learning rates around 1e-3 as an initial experiment, while SGD may require a different scale.

A second major issue is gradient explosion. If gradients become extremely large, parameter updates can become enormous. Eventually, the loss may become inf or NaN.

You can inspect gradient magnitudes directly:

for name, parameter in model.named_parameters():
    if parameter.grad is not None:
        print(
            name,
            parameter.grad.abs().mean().item(),
            parameter.grad.abs().max().item()
        )

This is useful because looking only at the loss does not tell you what is happening inside the optimization process. A model can have a rapidly increasing loss while some individual layers are receiving extremely large gradients.

Gradient clipping can help when large gradients are a known problem. PyTorch provides clip_grad_norm_() for this purpose:

torch.nn.utils.clip_grad_norm_(
    model.parameters(),
    max_norm=1.0
)

optimizer.step()

Gradient clipping does not fix every convergence problem. It limits the magnitude of the gradients before the optimizer updates the parameters, making it particularly useful in situations where occasional very large gradients destabilize training. The opposite problem is vanishing gradients. Here, gradients become extremely small as they propagate through the network. Early layers may then receive almost no useful signal, causing them to learn extremely slowly.


This problem historically appeared frequently in deep networks using saturating activation functions such as sigmoid and tanh. Modern architectures often use ReLU-family activations, residual connections, normalization layers, and carefully selected initialization methods to reduce these problems. Initialization itself can also influence convergence. Poorly scaled initial weights can cause activations or gradients to become too large or too small before training has even had a chance to make meaningful updates.

Another frequent problem has nothing to do with the optimizer: the target values and model outputs may be incompatible with the loss function.


For example, binary classification is commonly implemented with a single output logit and BCEWithLogitsLoss:

model = nn.Linear(input_size, 1)
criterion = nn.BCEWithLogitsLoss()

logits = model(x)
loss = criterion(logits, targets.float())

BCEWithLogitsLoss combines a sigmoid operation with binary cross-entropy in a numerically stable way. Applying a sigmoid manually before passing the output to this loss is unnecessary and can lead to an incorrect training setup. For multiclass classification, CrossEntropyLoss expects raw logits rather than probabilities:

model = nn.Linear(input_size, num_classes)
criterion = nn.CrossEntropyLoss()

logits = model(x)
loss = criterion(logits, targets)

Here, the target should normally contain class indices rather than one-hot encoded vectors. A mismatch between the expected target format and the supplied data can cause errors or, depending on the implementation, produce training behavior that does not match your expectations.


Data scaling is another major factor. Suppose one input feature ranges from 0 to 1, while another ranges from 0 to 1,000,000. The optimization landscape can become poorly conditioned, making it harder for gradient-based optimizers to find efficient parameter updates.


For numerical features, standardization is often useful:


x′ = ( x − μ ) / σ


The exact preprocessing strategy depends on the problem, but consistently scaled inputs generally make optimization easier.

Batch size can also affect convergence. Very small batches produce noisy gradient estimates, while very large batches produce more stable estimates but can change the optimization dynamics and require different learning-rate choices.

Another subtle issue is accidentally disabling gradient computation. Code executed inside torch.no_grad() does not build the normal autograd graph:

with torch.no_grad():
    predictions = model(x)
    loss = criterion(predictions, y)

loss.backward()

This will not work as a normal training step because gradients were disabled during the forward pass. torch.no_grad() is primarily intended for inference and evaluation.

Similarly, detaching tensors from the computational graph can prevent gradients from reaching earlier operations:

features = encoder(x)
features = features.detach()

output = classifier(features)
loss = criterion(output, y)
loss.backward()

In this example, the classifier can receive gradients, but the encoder will not receive gradients through features. This can be intentional in transfer learning, but accidental use of .detach() can make part of a model appear to stop learning.


How to Diagnose a PyTorch Model That Is Not Converging

When a model refuses to converge, changing the optimizer repeatedly is usually less useful than systematically checking the training pipeline. Start with the smallest possible experiment.

First, verify that the loss can decrease on a tiny subset of your data. Training on perhaps 10–100 samples is a useful debugging technique. A sufficiently flexible model should often be able to overfit a tiny dataset. If it cannot, there may be a problem with the model, loss function, labels, gradients, or training loop.

For example:

model.train()

for epoch in range(500):
    optimizer.zero_grad()

    output = model(x_small)
    loss = criterion(output, y_small)

    loss.backward()
    optimizer.step()

    if epoch % 50 == 0:
        print(f"Epoch {epoch}: {loss.item():.6f}")

If the loss remains almost unchanged even on a tiny dataset, investigate the fundamentals before increasing the dataset size or making the architecture more complicated.

Next, verify that parameters are actually receiving gradients:

for name, parameter in model.named_parameters():
    if parameter.requires_grad:
        if parameter.grad is None:
            print(f"{name}: no gradient")
        else:
            print(
                f"{name}: "
                f"{parameter.grad.abs().mean().item():.6e}"
            )

A None gradient and a gradient close to zero mean different things. None often indicates that the parameter was not connected to the computation that produced the loss, gradients were disabled, or the parameter does not participate in the current forward pass. A very small numerical gradient may instead indicate vanishing gradients or a model/data issue.

It is also worth checking whether the optimizer actually contains the parameters you expect:

optimizer = torch.optim.Adam(model.parameters(), lr=1e-3)

optimizer_parameters = sum(
    parameter.numel()
    for group in optimizer.param_groups
    for parameter in group["params"]
)

model_parameters = sum(
    parameter.numel()
    for parameter in model.parameters()
    if parameter.requires_grad
)

print("Model parameters:", model_parameters)
print("Optimizer parameters:", optimizer_parameters)

If these numbers differ unexpectedly, some trainable parameters may not have been passed to the optimizer. Another useful debugging step is to monitor the parameter values themselves. If weights remain unchanged, gradients or optimizer configuration may be the problem. If they change dramatically every iteration, the learning rate or gradient magnitude may need attention.


You should also check the data independently of the model. Inspect a few inputs and labels manually. Verify tensor shapes, data types, label ranges, missing values, normalization, and preprocessing. A surprisingly large number of apparent optimization failures originate in the dataset pipeline.


For example, if a classification dataset accidentally contains incorrect labels, the optimizer may be functioning perfectly while the model receives contradictory training signals.

Finally, distinguish between training loss and validation loss. A model that cannot reduce training loss has an optimization or data-fitting problem. A model whose training loss decreases but whose validation performance deteriorates has a different problem, commonly related to overfitting, distribution differences, or regularization. A practical debugging sequence is therefore:


1. Check the loss function and target format 2. Verify model outputs and tensor shapes 3. Confirm gradients exist 4. Check gradient magnitudes 5. Confirm optimizer parameters 6. Try a tiny dataset 7. Test a smaller learning rate 8. Test a larger learning rate 9. Check input scaling 10. Inspect training and validation behavior


The important point is that convergence is not controlled by a single setting. PyTorch Gradient Descent depends on the interaction between the data, model, loss function, gradients, optimizer, learning rate, initialization, and training loop.

Once those components are checked systematically, a non-converging model becomes much less mysterious. Instead of randomly changing architectures or optimizers, you can identify exactly where the learning process is breaking down and correct that part of the pipeline.


Conclusion

A PyTorch model that refuses to converge is not necessarily a sign that the architecture is wrong. Problems with the learning rate, gradient flow, loss function, input scaling, initialization, or training loop can all prevent effective optimization. The key is to diagnose these components systematically instead of changing several settings at once.

PyTorch Gradient Descent gives us the tools needed to inspect what is happening during training. Checking gradient values, verifying optimizer parameters, testing on a small dataset, and monitoring both training and validation loss can quickly narrow down the source of the problem. Once the underlying issue is identified, adjusting the optimizer or model becomes a deliberate choice rather than trial and error.

Understanding these fundamentals also makes it easier to work with more complex PyTorch models. When training becomes unstable or progress stalls, going back to the gradient calculation and parameter update process is often the fastest way to find the cause.

Get in touch for customized mentorship, research and freelance solutions tailored to your needs.

bottom of page