Data science for business [electronic book] : what you need to know about data mining and data-analytic thinking / Foster Provost and Tom Fawcett
By: Provost, Foster [author]
.
Contributor(s): Fawcett, Tom [author]
.
Material type:
BookPublisher: Sebastopol, CA. : O'Reilly Media Inc., [2013]Copyright date: ©2013Description: online resource (xviii, 384 pages) : color illustrations.Content type: text Media type: computer Carrier type: online resourceISBN: 9781449361327 (paperback); 1449361323 (paperback); 9781449374297 (e-book).Other title: What you need to know about data mining and data-analytic thinking [Cover title].Subject(s): Data mining| Item type | Current library | Call number | Status | Notes | |
|---|---|---|---|---|---|
| eBook | MTU Online eBook | 006.312 (Browse shelf(Opens below)) | Available | CIT Module DATA 6001 - Core Reading. CIT Module INFO 7016 - Core Reading. |
Enhanced descriptions from Syndetics:
Written by renowned data science experts Foster Provost and Tom Fawcett, Data Science for Business introduces the fundamental principles of data science, and walks you through the "data-analytic thinking" necessary for extracting useful knowledge and business value from the data you collect. This guide also helps you understand the many data-mining techniques in use today.
Based on an MBA course Provost has taught at New York University over the past ten years, Data Science for Business provides examples of real-world business problems to illustrate these principles. You'll not only learn how to improve communication between business stakeholders and data scientists, but also how participate intelligently in your company's data science projects. You'll also discover how to think data-analytically, and fully appreciate how data science methods can support business decision-making.
Understand how data science fits in your organization--and how you can use it for competitive advantage Treat data as a business asset that requires careful investment if you're to gain real value Approach business problems data-analytically, using the data-mining process to gather good data in the most appropriate way Learn general concepts for actually extracting knowledge from data Apply data science principles when interviewing data science job candidatesIncludes bibliographical references (pages 361-368) and index.
Provides an introduction to the fundamental principles of data science, walking the reader through the "data-analytic thinking" necessary for extracting useful knowledge and business value from collected data.
Electronic reproduction.: ProQuest LibCentral. Mode of access: World Wide Web.
Table of contents provided by Syndetics
- Preface (p. xi)
- 1 Introduction: Data-Analytic Thinking (p. 1)
- The Ubiquity of Data Opportunities (p. 1)
- Example: Hurricane Frances (p. 3)
- Example: Predicting Customer Churn (p. 4)
- Data Science, Engineering, and Data-Driven Decision Making (p. 4)
- Data Processing and "Big Data" (p. 7)
- From Big Data 1.0 to Big Data 2.0 (p. 8)
- Data and Data Science Capability as a Strategic Asset (p. 9)
- Data-Analytic Thinking (p. 12)
- This Book (p. 14)
- Data Mining and Data Science, Revisited (p. 14)
- Chemistry Is Not About Test Tubes: Data Science Versus the Work of the Data Scientist (p. 15)
- Summary (p. 16)
- 2 Business Problems and Data Science Solutions (p. 19)
- Fundamental concepts: A set of canonical data mining tasks; The data mining process; Supervised versus unsupervised data mining.
- From Business Problems to Data Mining Tasks (p. 19)
- Supervised Versus Unsupervised Methods (p. 24)
- Data Mining and Its Results (p. 25)
- The Data Mining Process (p. 26)
- Business Understanding (p. 28)
- Data Understanding (p. 28)
- Data Preparation (p. 30)
- Modeling (p. 31)
- Evaluation (p. 31)
- Deployment (p. 32)
- Implications for Managing the Data Science Team (p. 34)
- Other Analytics Techniques and Technologies (p. 35)
- Statistics (p. 35)
- Database Querying (p. 37)
- Data Warehousing (p. 38)
- Regression Analysis (p. 39)
- Machine Learning and Data Mining (p. 39)
- Answering Business Questions with These Techniques (p. 40)
- Summary (p. 41)
- 3 Introduction to Predictive Modeling: From Correlation to Supervised Segmentation. (p. 43)
- Fundamental concepts: Identifying informative attributes; Segmenting data by progressive attribute selection.
- Exemplary techniques: Finding correlations; Attribute/variable selection; Tree induction.
- Models, Induction, and Prediction (p. 44)
- Supervised Segmentation (p. 48)
- Selecting Informative Attributes (p. 49)
- Example: Attribute Selection with Information Gain (p. 56)
- Supervised Segmentation with Tree-Structured Models (p. 62)
- Visualizing Segmentations (p. 67)
- Trees as Sets of Rules (p. 71)
- Probability Estimation (p. 71)
- Example: Addressing the Churn Problem with Tree Induction (p. 73)
- Summary (p. 78)
- 4 Fitting a Model to Data (p. 81)
- Fundamental concepts: Finding "optimal" model parameters based on data; Choosing the goal for data mining; Objective functions; Loss functions.
- Exemplary techniques: Linear regression; Logistic regression; Support-vector machines.
- Classification via Mathematical Functions (p. 83)
- Linear Discriminant Functions (p. 85)
- Optimizing an Objective Function (p. 87)
- An Example of Mining a Linear Discriminant from Data (p. 88)
- Linear Discriminant Functions for Scoring and Ranking Instances (p. 90)
- Support Vector Machines, Briefly (p. 91)
- Regression via Mathematical Functions (p. 94)
- Class Probability Estimation and Logistic "Regression" (p. 96)
- Logistic Regression: Some Technical Details (p. 99)
- Example: Logistic Regression versus Tree Induction (p. 102)
- Nonlinear Functions, Support Vector Machines, and Neural Networks (p. 105)
- Summary (p. 108)
- 5 Overfitting and Its Avoidance (p. 111)
- Fundamental concepts: Generalization; Fitting and overfitting; Complexity control. Exemplary techniques: Cross-validation; Attribute selection; Tree pruning; Regularization.
- Generalization (p. 111)
- Overfitting (p. 113)
- Overfitting Examined (p. 113)
- Holdout Data and Fitting Graphs (p. 113)
- Overfitting in Tree Induction (p. 116)
- Overfitting in Mathematical Functions (p. 118)
- Example: Overfitting Linear Functions (p. 119)
- Example: Why Is Overfitting Bad? (p. 124)
- From Holdout Evaluation to Cross-Validation (p. 126)
- The Churn Dataset Revisited (p. 129)
- Learning Curves (p. 130)
- Overfitting Avoidance and Complexity Control (p. 133)
- Avoiding Overfitting with Tree Induction (p. 133)
- A General Method for Avoiding Overfitting (p. 134)
- Avoiding Overfitting for Parameter Optimization (p. 136)
- Summary (p. 140)
- 6 Similarity, Neighbors, and Clusters (p. 141)
- Fundamental concepts: Calculating similarity of objects described by data; Using similarity for prediction; Clustering as similarity-based segmentation.
- Exemplary techniques: Searching for similar entities; Nearest neighbor methods; Clustering methods; Distance metrics for calculating similarity.
- Similarity and Distance (p. 142)
- Nearest-Neighbor Reasoning (p. 144)
- Example: Whiskey Analytics (p. 144)
- Nearest Neighbors for Predictive Modeling (p. 146)
- How Many Neighbors and How Much Influence? (p. 149)
- Geometric Interpretation, Overfitting, and Complexity Control (p. 151)
- Issues with Nearest-Neighbor Methods (p. 154)
- Some Important Technical Details Relating to Similarities and Neighbors (p. 157)
- Heterogeneous Attributes (p. 157)
- Other Distance Functions (p. 158)
- Combining Functions: Calculating Scores from Neighbors (p. 161)
- Clustering (p. 163)
- Example: Whiskey Analytics Revisited (p. 163)
- Hierarchical Clustering (p. 164)
- Nearest Neighbors Revisited: Clustering Around Centroids (p. 169)
- Example: Clustering Business News Stories (p. 174)
- Understanding the Results of Clustering (p. 177)
- Using Supervised Learning to Generate Cluster Descriptions (p. 179)
- Stepping Back: Solving a Business Problem Versus Data Exploration (p. 182)
- Summary (p. 184)
- 7 Decision Analytic Thinking I: What Is a Good Model? (p. 187)
- Fundamental concepts: Careful consideration of what is desired from data science results; Expected value as a key evaluation framework; Consideration of appropriate comparative baselines.
- Exemplary techniques: Various evaluation metrics; Estimating costs and benefits; Calculating expected profit; Creating baseline methods for comparison.
- Evaluating Classifiers (p. 188)
- Plain Accuracy and Its Problems (p. 189)
- The Confusion Matrix (p. 189)
- Problems with Unbalanced Classes (p. 190)
- Problems with Unequal Costs and Benefits (p. 193)
- Generalizing Beyond Classification (p. 193)
- A Key Analytical Framework: Expected Value (p. 194)
- Using Expected Value to Frame Classifier Use (p. 195)
- Using Expected Value to Frame Classifier Evaluation (p. 196)
- Evaluation, Baseline Performance, and Implications for Investments in Data (p. 204)
- Summary (p. 207)
- 8 Visualizing Model Performance (p. 209)
- Fundamental concepts: Visualization of model performance under various kinds of uncertainty; Further consideration of what is desired from data mining results.
- Exemplary techniques: Profit curves; Cumulative response curves; Lift curves; ROC curves.
- Ranking Instead of Classifying (p. 209)
- Profit Curves (p. 212)
- ROC Graphs and Curves (p. 214)
- The Area Under the ROC Curve (AUC) (p. 219)
- Cumulative Response and Lift Curves (p. 219)
- Example: Performance Analytics for Churn Modeling (p. 223)
- Summary (p. 231)
- 9 Evidence and Probabilities (p. 233)
- Fundamental concepts: Explicit evidence combination with Bayes' Rule; Probabilistic reasoning via assumptions of conditional independence.
- Exemplary techniques: Naive Bayes classification; Evidence lift.
- Example: Targeting Online Consumers With Advertisements (p. 233)
- Combining Evidence Probabilistically (p. 235)
- Joint Probability and Independence (p. 236)
- Bayes' Rule (p. 237)
- Applying Bayes' Rule to Data Science (p. 239)
- Conditional Independence and Naive Bayes (p. 240)
- Advantages and Disadvantages of Naive Bayes (p. 242)
- A Model of Evidence "Lift" (p. 244)
- Example: Evidence Lifts from Facebook "Likes" (p. 245)
- Evidence in Action: Targeting Consumers with Ads (p. 247)
- Summary (p. 247)
- 10 Representing and Mining Text (p. 249)
- Fundamental concepts: The importance of constructing mining-friendly data representations; Representation of text for data mining.
- Exemplary techniques: Bag of words representation; TFIDF calculation; N-grams; Stemming; Named entity extraction; Topic models.
- Why Text Is Important (p. 250)
- Why Text Is Difficult (p. 250)
- Representation (p. 251)
- Bag of Words (p. 252)
- Term Frequency (p. 252)
- Measuring Sparseness: Inverse Document Frequency (p. 254)
- Combining Them: TFIDF (p. 256)
- Example: Jazz Musicians (p. 256)
- The Relationship of IDF to Entropy (p. 261)
- Beyond Bag of Words (p. 263)
- N-gram Sequences (p. 263)
- Named Entity Extraction (p. 264)
- Topic Models (p. 264)
- Example: Mining News Stories to Predict Stock Price Movement (p. 266)
- The Task (p. 266)
- The Data (p. 268)
- Data Preprocessing (p. 271)
- Results (p. 271)
- Summary (p. 275)
- 11 Decision Analytic Thinking II: Toward Analytical Engineering (p. 277)
- Fundamental concept: Solving business problems with data science starts with analytical engineering: designing an analytical solution, based on the data, tools, and techniques available.
- Exemplary technique: Expected value as a framework for data science solution design.
- Targeting the Best Prospects for a Charity Mailing (p. 278)
- The Expected Value Framework: Decomposing the Business Problem and Recomposing the Solution Pieces (p. 278)
- A Brief Digression on Selection Bias (p. 280)
- Our Churn Example Revisited with Even More Sophistication (p. 281)
- The Expected Value Framework: Structuring a More Complicated Business Problem (p. 281)
- Assessing the Influence of the Incentive (p. 283)
- From an Expected Value Decomposition to a Data Science Solution (p. 284)
- Summary (p. 287)
- 12 Other Data Science Tasks and Techniques (p. 289)
- Fundamental concepts: Our fundamental concepts as the basis of many common data science techniques; The importance of familiarity with the building blocks of data science.
- Exemplary techniques: Association and co - occurrences; Behavior profiling; Link prediction; Data reduction; Latent information mining; Movie recommendation; Bias-variance decomposition of error; Ensembles of models; Causal reasoning from data.
- Co-occurrences and Associations: Finding Items That Go Together (p. 290)
- Measuring Surprise: Lift and Leverage (p. 291)
- Example: Beer and Lottery Tickets (p. 292)
- Associations Among Facebook Likes (p. 293)
- Profiling: Finding Typical Behavior (p. 296)
- Link Prediction and Social Recommendation (p. 301)
- Data Reduction, Latent Information, and Movie Recommendation (p. 302)
- Bias, Variance, and Ensemble Methods (p. 306)
- Data-Driven Causal Explanation and a Viral Marketing Example (p. 309)
- Summary (p. 310)
- 13 Data Science and Business Strategy (p. 313)
- Fundamental concepts: Our principles as the basis of success for a data-driven business; Acquiring and sustaining competitive advantage via data science; The importance of careful curation of data science capability.
- Thinking Data-Analytically, Redux (p. 313)
- Achieving Competitive Advantage with Data Science (p. 315)
- Sustaining Competitive Advantage with Data Science (p. 316)
- Formidable Historical Advantage (p. 317)
- Unique Intellectual Property (p. 317)
- Unique Intangible Collateral Assets (p. 318)
- Superior Data Scientists (p. 318)
- Superior Data Science Management (p. 320)
- Attracting and Nurturing Data Scientists and Their Teams (p. 321)
- Examine Data Science Case Studies (p. 323)
- Be Ready to Accept Creative Ideas from Any Source (p. 324)
- Be Ready to Evaluate Proposals for Data Science Projects (p. 324)
- Example Data Mining Proposal (p. 325)
- Flaws in the Big Red Proposal (p. 326)
- A Firm's Data Science Maturity (p. 327)
- 14 Conclusion (p. 331)
- The Fundamental Concepts of Data Science (p. 331)
- Applying Our Fundamental Concepts to a New Problem: Mining Mobile Device Data (p. 334)
- Changing the Way We Think about Solutions to Business Problems (p. 337)
- What Data Can't Do: Humans in the Loop, Revisited (p. 338)
- Privacy, Ethics, and Mining Data About Individuals (p. 341)
- Is There More to Data Science? (p. 342)
- Final Example: From Crowd-Sourcing to Cloud-Sourcing (p. 343)
- Final Words (p. 344)
- A Proposal Review Guide (p. 347)
- B Another Sample Proposal (p. 351)
- Glossary (p. 355)
- Bibliography (p. 359)
- Index (p. 367)
Author notes provided by Syndetics
Foster Provost is Professor and NEC Faculty Fellow at the NYU Stern School of Business, where he teaches in the MBA, Business Analytics, and Data Science programs. Former Editor-in-Chief for the journal Machine Learning, Professor Provost has co-founded several successful companies focusing on data science for marketing.
Tom Fawcett holds a Ph.D. in machine learning and has worked in industry R&D for more than two decades for companies such as GTE Laboratories, NYNEX/Verizon Labs, and HP Labs. His published work has become standard reading in data science both on methodology (evaluating data mining results) and on applications (fraud detection and spam filtering).