Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

4 Commits

Folders and files

Repository files navigation

Shallow NN vs Logistic Regression

Does adding one hidden layer actually buy you anything over a linear model?

This repo trains both on the same data, written from scratch in NumPy, and reports the difference. The short answer is it depends entirely on whether the boundary is non-linear — and the first hidden unit on its own buys you nothing at all.

Accuracy and AUC against number of hidden units

The result

MAGIC Gamma Telescope (19,020 events, 10 features), stratified 80/20 split, seed 0. AUC is reported alongside accuracy because UCI specifies AUC for this dataset — a hadron misread as a gamma is the expensive error, so plain accuracy understates the difference.

Model Accuracy AUC
logistic regression (scikit-learn) 0.7905 0.8439
shallow NN, 1 hidden unit 0.7900 0.8436
shallow NN, 2 hidden units 0.8431 0.8960
shallow NN, 4 hidden units 0.8573 0.9064
shallow NN, 8 hidden units 0.8672 0.9185
shallow NN, 16 hidden units 0.8693 0.9226
shallow NN, 32 hidden units 0.8680 0.9244
shallow NN, 64 hidden units 0.8680 0.9236

+0.0805 AUC over logistic regression, and it plateaus around 16–32 units.

Why the first hidden unit does nothing

With one hidden unit the network computes sigmoid(w2 · relu(w1·x + b1) + b2). Since both activations are monotonic, the decision boundary is still the set where w1·x + b1 crosses a constant — a straight line, exactly the family of boundaries logistic regression can draw.

The table shows this rather than asserting it: 1 hidden unit scores 0.8436 AUC against logistic regression's 0.8439. A difference of 0.0003, i.e. nothing. One unit is a linear model wearing a neural network's clothes.

The non-linearity only appears from the second unit onward, and that is where the accuracy comes from.

Training loss per epoch

The loss curves tell the same story from the optimisation side: more hidden units reach a lower training loss, and the ordering is monotonic until the plateau.

Why this matters for the course material

This is week 3 of Andrew Ng's Neural Networks and Deep Learning made measurable. The relevant intuitions, and where they show up here:

  • A hidden layer learns features, it does not just add parameters. Units compose into a piecewise non-linear boundary; one unit cannot.
  • Weight initialisation must break symmetry. nn.init_params uses small random weights — zero weights would make every hidden unit identical and the network would never learn anything but a linear model.
  • Capacity controls boundary complexity. The capacity curve is the practical version of that claim.

nn.py implements forward propagation, binary cross-entropy, backpropagation, and gradient descent directly — no scikit-learn for the network itself. The library is used only for the logistic regression baseline and the train/test split, so the comparison is not a strawman.

The data

MAGIC Gamma Telescope, UCI Machine Learning Repository (dataset 159, CC BY 4.0). Ground-based Cherenkov telescope events described by the Hillas parameters of their shower image; the task is separating gamma-ray showers from cosmic-ray hadron showers.

Two things worth stating plainly, both documented in data/README.md:

  • It is Monte Carlo simulated, not observed. Generated with the Corsika shower simulation to model the detector. It is a standard benchmark, but it is not raw telescope output.
  • The source file is grouped by class — the first half is all gamma. Rows are independent events, so stratified cross-validation is valid; a chronological split would not be. The pipeline always shuffles and stratifies.

Reproduce

python -m venv .venv
.venv\Scripts\activate
pip install -r requirements.txt
python run.py

Runs on CPU in well under a minute and writes both figures to images/.

What else I tried, and why it is not here

Choosing the dataset took longer than building the model, and the failures are more instructive than the success. Four datasets were measured and rejected:

Dataset n Log. reg. Best NN Gap Why rejected
EEG Eye State 14,980 0.5975 0.8222 +0.2248 temporal leakage
Diabetic Retinopathy Debrecen 1,151 0.7193 0.7402 +0.0209 gap inside fold noise
Breast Cancer Wisconsin 569 0.9580 0.9580 +0.0000 already linearly separable
Banknote 1,372 0.9811 0.9993 +0.0182 already linearly separable

The EEG result is the one worth reading. It looked like the strongest finding in the whole project — a 22-point gap. It was an artefact.

EEG is a time series: consecutive rows are 99.2% identical. A random train/test split therefore puts near-duplicate samples on both sides of the split, and the network memorises them. Evaluated honestly, on a chronological split:

Evaluation Logistic regression NN (16 units)
Random 5-fold CV 0.5975 0.8222
Chronological 70/30 0.3331 0.4355
Chronological 5-fold 0.4144 0.4137
Majority-class baseline 0.5512

The 22-point advantage becomes a 0.07-point tie, and both models land below the majority-class baseline. The entire result was leakage.

The lesson generalised: on small tabular data, logistic regression is very hard to beat. The clear win only appeared on a dataset with genuine non-linearity and enough rows (19,020) for the hidden layer to find it.

Files

File What it is
nn.py the shallow network: init, forward, loss, backprop, train
run.py trains both models, prints the table, saves the figures
data/magic_gamma.csv the dataset, vendored
data/README.md provenance, license, checksum, caveats
plan.md, task.md planning notes, kept for transparency

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages