Syllabus
Unit 1
Introduction to Data Mining Introduction, what is Data Mining? Concepts of Data mining, Technologies Used, Data Mining Process, KDD Process Model, CRISP DM, Mining on various kinds of data, Applications of Data Mining, Challenges of Data Mining.
Unit 2
Data Understanding and Preparation Introduction, Reading data from various sources, Data visualization, Distributions and summary statistics, Relationships among variables, Extent of Missing Data. Segmentation, Outlier detection, Automated Data Preparation, combining data files, Aggregate Data, Duplicate Removal, Sampling DATA, Data Caching, Partitioning data, Missing Values.
Unit 3
Model development & techniques Data Partitioning, Model selection, Model Development Techniques, Neural networks, Decision trees, Logistic regression, Discriminant analysis, Support vector machine, Bayesian Networks, Linear Regression, Cox Regression, Association rules.
Unit 4
Model Evaluation and Deployment Introduction, Model Validation, Rule Induction Using CHAID, Automating Models for Categorical and Continuous targets, Comparing and Combining Models, Evaluation Charts for Model Comparison, Meta Level Modeling, Deploying Model, Assessing Model Performance, Updating a Model.
Text Books / References
Text Books:
- Eric Siegel. (2016) Predictive Analytics: The Power to Predict Who Will Click, Buy, Lie, or Die. John Wiley & Sons, Hoboken, New Jersey.
- Dean Abbott. (2014). Applied Predictive Analytics: Principles and Techniques for the Professional Data Analyst. Wiley.
References:
- David L Olson (2016) Data Mining Models (2nd ed.), Business Expert Press.
- Jeffrey T. Prince and Amarnath Bose. (2020) A Predictive Analytics for Business Strategy – Reasoning from Data to Actionable Knowledge, Mc Graw Hill.
Introduction
: This course aims to equip students with skills to analyse historic data, build machine learning models (regression, classification, forecasting), and apply these to solve real-world business problems. Objectives focus on data pre-processing, model evaluation, and utilizing software tools to predict future outcomes.
Objectives and Outcomes
Course Outcomes: After successful completion of the course, students will be able to
- CO1: Extract key information from the data by using data mining techniques
- CO2: Fetch data from various sources and apply data pre-processing techniques
- CO3: Choose appropriate model for data analysis
- CO4: Evaluate and implement classification algorithms
- CO5: Evaluate the accuracy and performance of different data mining models and algorithms
CO-PO Mapping:
| |
PO1 |
PO2 |
PO3 |
PO4 |
PO5 |
PO6 |
PO7 |
PO8 |
PO9 |
PO10 |
PO11 |
PO12 |
| CO1 |
2 |
3 |
3 |
2 |
3 |
3 |
|
|
|
|
2 |
|
| CO2 |
2 |
3 |
3 |
2 |
3 |
3 |
|
|
|
|
2 |
|
| CO3 |
2 |
3 |
1 |
2 |
1 |
3 |
|
|
|
|
2 |
|
| CO4 |
2 |
3 |
3 |
2 |
3 |
3 |
|
|
|
|
2 |
|
| CO5 |
2 |
3 |
3 |
2 |
3 |
3 |
|
|
|
|
2 |
|