The book is based on Stanford Computer Science course CS246: Mining Massive Datasets (and CS345A: Data Mining).