Without a Nonlinearity the Whole Network Collapses to One Layer
I remember sitting in a windowless lab three years ago, watching a training loss curve flatten into a perfectly straight, useless line. I had spent forty-eight hours tuning hyperparameters, only to realize I had blindly defaulted to a Sigmoid function…
Read More








