Back close

Course Detail

Course Name Mining of Massive Datasets
Course Code 26CSC336
Program 5 Year Integrated M.Sc in Data Science
Credits 3
Campus Coimbatore

Syllabus

Syllabus

Basics of Data Mining – Computational Approaches – Statistical Limits on Data Mining – Bonferroni’s Principle – MapReduce – Distributed File Systems. MapReduce. Algorithms Using MapReduce. Extensions to MapReduce. Finding Similar Items – Applications of Near Neighbor Search – Shingling of Documents – Similarity-Preserving Summaries of Sets – Locality-Sensitive Hashing for Documents – Distance Measures.

Mining Data Streams: The Stream Data Model – Sampling Data in a Stream – Filtering Streams. Link Analysis: PageRank – Efficient Computation of PageRank – Topic-Sensitive PageRank – Link Spam.

Frequent Itemsets: The Market-Basket Model – Market Baskets and the A-Priori Algorithm – Handling Larger Datasets in Main Memory.

Clustering: Introduction to Clustering Techniques – Hierarchical Clustering – K-means Algorithms – CURE algorithm. Recommendation Systems: A Model for Recommendation Systems – Content-Based Recommendations – Collaborative Filtering – Dimensionality Reduction.

Mining Social Network Graphs: Social Networks as Graphs – Clustering of Social-Network Graphs – Direct Discovery of Communities – Partitioning of Graphs – Finding Overlapping Communities – Simrank.

Dimensionality Reduction: Eigenvalues and Eigenvectors of Symmetric Matrices- Principal-Component Analysis – Singular-Value Decomposition. Large-Scale Machine Learning – Machine-Learning Model – Perceptrons – Support-Vector Machines.

Text Books / References

Text / References Book

  1. Jure Leskovec, Anand Rajaraman, Jeffrey David Ullman, Mining of Massive Datasets, Cambridge University Press, 2014.
  2. Tom White, Hadoop: The Definitive Guide: Storage and Analysis at Internet Scale O’Reilly Media; 4th edition, 2015.

Introduction

This course introduces the principles and techniques of data mining for large-scale and distributed data environments. It begins with foundational concepts, computational approaches, and statistical limitations inherent in extracting knowledge from massive datasets. The course explores distributed data processing frameworks such as MapReduce and Distributed File Systems, including algorithm design and extensions for scalable analytics. It further covers mining data streams using stream data models, sampling techniques, and filtering methods for real-time processing. Key data mining applications such as link analysis, frequent itemset mining, clustering, recommendation systems, web advertising, and social network graph mining are discussed. Advanced topics including dimensionality reduction and large-scale machine learning provide insights into handling high-dimensional and big data scenarios. The course equips students with theoretical understanding and practical strategies for scalable data mining in modern distributed systems.

Objectives and Outcomes

Course Outcomes: After successful completion of the course, students will be able to

  • CO1: Explain the fundamental principles of data mining, including computational approaches, statistical limitations, and challenges in large-scale data processing environments.
  • CO2: Design and implement distributed data processing solutions using MapReduce and Distributed File Systems, including the development of scalable algorithms and extensions for big data applications.
  • CO3: Apply stream mining techniques under the stream data model, including sampling, filtering, and real-time data processing methods.
  • CO4: Develop and evaluate data mining solutions for real-world applications such as link analysis, frequent itemset mining, clustering, recommendation systems, social network analysis, dimensionality reduction, and large-scale machine learning.

CO-PO Mapping

  PO1 PO2 PO3 PO4 PO5 PO6 PO7 PO8 PO9 PO10 PO11 PO12
CO1 2 2 2 2 2 2         1 1
CO2 3 3 2 2 2 2         1 1
CO3 2 2 3 2 2 2         1 1
CO4 3 3 3 2 2 2         1 1

DISCLAIMER: The appearance of external links on this web site does not constitute endorsement by the School of Biotechnology/Amrita Vishwa Vidyapeetham or the information, products or services contained therein. For other than authorized activities, the Amrita Vishwa Vidyapeetham does not exercise any editorial control over the information you may find at these locations. These links are provided consistent with the stated purpose of this web site.

Admissions Apply Now