Does adding one hidden layer actually buy you anything over a linear model?
This repo trains both on the same data, written from scratch in NumPy, and reports the difference. The short answer is it depends entirely on whether the boundary is non-linear — and the first hidden unit on its own buys you nothing at all.
MAGIC Gamma Telescope (19,020 events, 10 features), stratified 80/20 split, seed 0. AUC is reported alongside accuracy because UCI specifies AUC for this dataset — a hadron misread as a gamma is the expensive error, so plain accuracy understates the difference.
| Model | Accuracy | AUC |
|---|---|---|
| logistic regression (scikit-learn) | 0.7905 | 0.8439 |
| shallow NN, 1 hidden unit | 0.7900 | 0.8436 |
| shallow NN, 2 hidden units | 0.8431 | 0.8960 |
| shallow NN, 4 hidden units | 0.8573 | 0.9064 |
| shallow NN, 8 hidden units | 0.8672 | 0.9185 |
| shallow NN, 16 hidden units | 0.8693 | 0.9226 |
| shallow NN, 32 hidden units | 0.8680 | 0.9244 |
| shallow NN, 64 hidden units | 0.8680 | 0.9236 |
+0.0805 AUC over logistic regression, and it plateaus around 16–32 units.
Why the first hidden unit does nothing
With one hidden unit the network computes sigmoid(w2 · relu(w1·x + b1) + b2).
Since both activations are monotonic, the decision boundary is still the set
where w1·x + b1 crosses a constant — a straight line, exactly the family
of boundaries logistic regression can draw.
The table shows this rather than asserting it: 1 hidden unit scores 0.8436 AUC against logistic regression's 0.8439. A difference of 0.0003, i.e. nothing. One unit is a linear model wearing a neural network's clothes.
The non-linearity only appears from the second unit onward, and that is where the accuracy comes from.
The loss curves tell the same story from the optimisation side: more hidden units reach a lower training loss, and the ordering is monotonic until the plateau.
This is week 3 of Andrew Ng's Neural Networks and Deep Learning made measurable. The relevant intuitions, and where they show up here:
- A hidden layer learns features, it does not just add parameters. Units compose into a piecewise non-linear boundary; one unit cannot.
- Weight initialisation must break symmetry.
nn.init_paramsuses small random weights — zero weights would make every hidden unit identical and the network would never learn anything but a linear model. - Capacity controls boundary complexity. The capacity curve is the practical version of that claim.
nn.py implements forward propagation, binary cross-entropy,
backpropagation, and gradient descent directly — no scikit-learn for the
network itself. The library is used only for the logistic regression baseline
and the train/test split, so the comparison is not a strawman.
MAGIC Gamma Telescope, UCI Machine Learning Repository (dataset 159, CC BY 4.0). Ground-based Cherenkov telescope events described by the Hillas parameters of their shower image; the task is separating gamma-ray showers from cosmic-ray hadron showers.
Two things worth stating plainly, both documented in data/README.md:
- It is Monte Carlo simulated, not observed. Generated with the Corsika shower simulation to model the detector. It is a standard benchmark, but it is not raw telescope output.
- The source file is grouped by class — the first half is all gamma. Rows are independent events, so stratified cross-validation is valid; a chronological split would not be. The pipeline always shuffles and stratifies.
python -m venv .venv
.venv\Scripts\activate
pip install -r requirements.txt
python run.pyRuns on CPU in well under a minute and writes both figures to images/.
Choosing the dataset took longer than building the model, and the failures are more instructive than the success. Four datasets were measured and rejected:
| Dataset | n | Log. reg. | Best NN | Gap | Why rejected |
|---|---|---|---|---|---|
| EEG Eye State | 14,980 | 0.5975 | 0.8222 | +0.2248 | temporal leakage |
| Diabetic Retinopathy Debrecen | 1,151 | 0.7193 | 0.7402 | +0.0209 | gap inside fold noise |
| Breast Cancer Wisconsin | 569 | 0.9580 | 0.9580 | +0.0000 | already linearly separable |
| Banknote | 1,372 | 0.9811 | 0.9993 | +0.0182 | already linearly separable |
The EEG result is the one worth reading. It looked like the strongest finding in the whole project — a 22-point gap. It was an artefact.
EEG is a time series: consecutive rows are 99.2% identical. A random train/test split therefore puts near-duplicate samples on both sides of the split, and the network memorises them. Evaluated honestly, on a chronological split:
| Evaluation | Logistic regression | NN (16 units) |
|---|---|---|
| Random 5-fold CV | 0.5975 | 0.8222 |
| Chronological 70/30 | 0.3331 | 0.4355 |
| Chronological 5-fold | 0.4144 | 0.4137 |
| Majority-class baseline | 0.5512 |
The 22-point advantage becomes a 0.07-point tie, and both models land below the majority-class baseline. The entire result was leakage.
The lesson generalised: on small tabular data, logistic regression is very hard to beat. The clear win only appeared on a dataset with genuine non-linearity and enough rows (19,020) for the hidden layer to find it.
| File | What it is |
|---|---|
nn.py |
the shallow network: init, forward, loss, backprop, train |
run.py |
trains both models, prints the table, saves the figures |
data/magic_gamma.csv |
the dataset, vendored |
data/README.md |
provenance, license, checksum, caveats |
plan.md, task.md |
planning notes, kept for transparency |

