The Sigmoid function maps inputs to a value between 0 and 1, often interpreted as a probability in binary classification. It can saturate for large positive or negative inputs, leading to slow gradient updates. Nevertheless, it remains a common choice in logistic regression and output layers.