Aegis: A Root Cause Engine for Distributed Telemetry
The Anatomy of a Root Cause Engine: Correlating Hidden Patterns in Distributed Telemetry
The Anatomy of a Root Cause Engine: Correlating Hidden Patterns in Distributed Telemetry
Information gain is a very important concept in information theory. In the most simplest case, it is the reduction in entropy. If you are not familiar with entropy, check out my entropy post. Information gain is used in decision trees. A decision tree is a classification algorithm. Decision tree splits the given feature from specific points with specific questions by branching. By doing that, it tries to maximize information gain. In other words, it tries to minimize the entropy in each branch. To sum up, information gain tells us the reduction in entropy in each split.
Entropy is one of the key concepts of information theory, data science and machine learning. Understanding the math behind it is crucial for designing solid machine learning pipelines. Therefore, in this post, I try to explain the entropy as simple as possible.