MTU Library Catalogue

Syndetics cover image
Image from Syndetics

Applied data mining : statistical methods for business and industry / Paolo Giudici.

By: Giudici, Paolo.
Material type: materialTypeLabelBookPublisher: New York : J. Wiley, 2003Description: xii, 364 p. : ill. : 24 cm.ISBN: 0470846798 (pbk.); 9780470846797; 9780470846780.Subject(s): Data mining | Business -- Data processing | Commercial statistics | Business -- Data processing -- Statistical methods | Business and Management | Calculus & mathematical analysis | Computer hardware | Business applications | Business mathematics & systems | DatabasesDDC classification: 658.056312 GIU Summary: Data mining can be defined as the process of selection, exploration and modelling of large databases, in order to discover models and patterns. This text offers business students and industry professionals an accessible introduction to this increasingly important field.
Holdings
Item type Current library Call number Copy number Status Barcode
General lending MTU Kerry North Campus Library First Floor Main 658.056312 GIU (Browse shelf(Opens below)) 1 Available 38888000547889
Total holds: 0

Enhanced descriptions from Syndetics:

Data mining can be defined as the process of selection, exploration and modelling of large databases, in order to discover models and patterns. The increasing availability of data in the current information society has led to the need for valid tools for its modelling and analysis. Data mining and applied statistical methods are the appropriate tools to extract such knowledge from data. Applications occur in many different fields, including statistics, computer science, machine learning, economics, marketing and finance.

This book is the first to describe applied data mining methods in a consistent statistical framework, and then show how they can be applied in practice. All the methods described are either computational, or of a statistical modelling nature. Complex probabilistic models and mathematical tools are not used, so the book is accessible to a wide audience of students and industry professionals. The second half of the book consists of nine case studies, taken from the author's own work in industry, that demonstrate how the methods described can be applied to real problems.

Provides a solid introduction to applied data mining methods in a consistent statistical framework Includes coverage of classical, multivariate and Bayesian statistical methodology Includes many recent developments such as web mining, sequential Bayesian analysis and memory based reasoning Each statistical method described is illustrated with real life applications Features a number of detailed case studies based on applied projects within industry Incorporates discussion on software used in data mining, with particular emphasis on SAS Supported by a website featuring data sets, software and additional material Includes an extensive bibliography and pointers to further reading within the text Author has many years experience teaching introductory and multivariate statistics and data mining, and working on applied projects within industry

A valuable resource for advanced undergraduate and graduate students of applied statistics, data mining, computer science and economics, as well as for professionals working in industry on projects involving large volumes of data - such as in marketing or financial risk management.

Includes bibliographical references (p. [353]-356) and index.

Data mining can be defined as the process of selection, exploration and modelling of large databases, in order to discover models and patterns. This text offers business students and industry professionals an accessible introduction to this increasingly important field.

Table of contents provided by Syndetics

  • Preface (p. xi)
  • 1 Introduction (p. 1)
  • 1.1 What is data mining? (p. 1)
  • 1.1.1 Data mining and computing (p. 3)
  • 1.1.2 Data mining and statistics (p. 5)
  • 1.2 The data mining process (p. 6)
  • 1.3 Software for data mining (p. 11)
  • 1.4 Organisation of the book (p. 12)
  • 1.4.1 Chapters 2 to 6: methodology (p. 13)
  • 1.4.2 Chapters 7 to 12: business cases (p. 13)
  • 1.5 Further reading (p. 14)
  • Part I Methodology (p. 17)
  • 2 Organisation of the data (p. 19)
  • 2.1 From the data warehouse to the data marts (p. 20)
  • 2.1.1 The data warehouse (p. 20)
  • 2.1.2 The data webhouse (p. 21)
  • 2.1.3 Data marts (p. 22)
  • 2.2 Classification of the data (p. 22)
  • 2.3 The data matrix (p. 23)
  • 2.3.1 Binarisation of the data matrix (p. 25)
  • 2.4 Frequency distributions (p. 25)
  • 2.4.1 Univariate distributions (p. 26)
  • 2.4.2 Multivariate distributions (p. 27)
  • 2.5 Transformation of the data (p. 29)
  • 2.6 Other data structures (p. 30)
  • 2.7 Further reading (p. 31)
  • 3 Exploratory data analysis (p. 33)
  • 3.1 Univariate exploratory analysis (p. 34)
  • 3.1.1 Measures of location (p. 35)
  • 3.1.2 Measures of variability (p. 37)
  • 3.1.3 Measures of heterogeneity (p. 37)
  • 3.1.4 Measures of concentration (p. 39)
  • 3.1.5 Measures of asymmetry (p. 41)
  • 3.1.6 Measures of kurtosis (p. 43)
  • 3.2 Bivariate exploratory analysis (p. 45)
  • 3.3 Multivariate exploratory analysis of quantitative data (p. 49)
  • 3.4 Multivariate exploratory analysis of qualitative data (p. 51)
  • 3.4.1 Independence and association (p. 53)
  • 3.4.2 Distance measures (p. 54)
  • 3.4.3 Dependency measures (p. 56)
  • 3.4.4 Model-based measures (p. 58)
  • 3.5 Reduction of dimensionality (p. 61)
  • 3.5.1 Interpretation of the principal components (p. 63)
  • 3.5.2 Application of the principal components (p. 65)
  • 3.6 Further reading (p. 66)
  • 4 Computational data mining (p. 69)
  • 4.1 Measures of distance (p. 70)
  • 4.1.1 Euclidean distance (p. 71)
  • 4.1.2 Similarity measures (p. 72)
  • 4.1.3 Multidimensional scaling (p. 74)
  • 4.2 Cluster analysis (p. 75)
  • 4.2.1 Hierarchical methods (p. 77)
  • 4.2.2 Evaluation of hierarchical methods (p. 81)
  • 4.2.3 Non-hierarchical methods (p. 83)
  • 4.3 Linear regression (p. 85)
  • 4.3.1 Bivariate linear regression (p. 85)
  • 4.3.2 Properties of the residuals (p. 88)
  • 4.3.3 Goodness of fit (p. 90)
  • 4.3.4 Multiple linear regression (p. 91)
  • 4.4 Logistic regression (p. 96)
  • 4.4.1 Interpretation of logistic regression (p. 97)
  • 4.4.2 Discriminant analysis (p. 98)
  • 4.5 Tree models (p. 100)
  • 4.5.1 Division criteria (p. 103)
  • 4.5.2 Pruning (p. 105)
  • 4.6 Neural networks (p. 107)
  • 4.6.1 Architecture of a neural network (p. 109)
  • 4.6.2 The multilayer perceptron (p. 111)
  • 4.6.3 Kohonen networks (p. 117)
  • 4.7 Nearest-neighbour models (p. 119)
  • 4.8 Local models (p. 121)
  • 4.8.1 Association rules (p. 121)
  • 4.8.2 Retrieval by content (p. 126)
  • 4.9 Further reading (p. 127)
  • 5 Statistical data mining (p. 129)
  • 5.1 Uncertainty measures and inference (p. 129)
  • 5.1.1 Probability (p. 130)
  • 5.1.2 Statistical models (p. 132)
  • 5.1.3 Statistical inference (p. 137)
  • 5.2 Non-parametric modelling (p. 143)
  • 5.3 The normal linear model (p. 146)
  • 5.3.1 Main inferential results (p. 147)
  • 5.3.2 Application (p. 150)
  • 5.4 Generalised linear models (p. 154)
  • 5.4.1 The exponential family (p. 155)
  • 5.4.2 Definition of generalised linear models (p. 157)
  • 5.4.3 The logistic regression model (p. 163)
  • 5.4.4 Application (p. 164)
  • 5.5 Log-linear models (p. 167)
  • 5.5.1 Construction of a log-linear model (p. 167)
  • 5.5.2 Interpretation of a log-linear model (p. 169)
  • 5.5.3 Graphical log-linear models (p. 171)
  • 5.5.4 Log-linear model comparison (p. 174)
  • 5.5.5 Application (p. 175)
  • 5.6 Graphical models (p. 177)
  • 5.6.1 Symmetric graphical models (p. 178)
  • 5.6.2 Recursive graphical models (p. 182)
  • 5.6.3 Graphical models versus neural networks (p. 184)
  • 5.7 Further reading (p. 185)
  • 6 Evaluation of data mining methods (p. 187)
  • 6.1 Criteria based on statistical tests (p. 188)
  • 6.1.1 Distance between statistical models (p. 188)
  • 6.1.2 Discrepancy of a statistical model (p. 190)
  • 6.1.3 The Kullback--Leibler discrepancy (p. 192)
  • 6.2 Criteria based on scoring functions (p. 193)
  • 6.3 Bayesian criteria (p. 195)
  • 6.4 Computational criteria (p. 197)
  • 6.5 Criteria based on loss functions (p. 200)
  • 6.6 Further reading (p. 204)
  • Part II Business cases (p. 207)
  • 7 Market basket analysis (p. 209)
  • 7.1 Objectives of the analysis (p. 209)
  • 7.2 Description of the data (p. 210)
  • 7.3 Exploratory data analysis (p. 212)
  • 7.4 Model building (p. 215)
  • 7.4.1 Log-linear models (p. 215)
  • 7.4.2 Association rules (p. 218)
  • 7.5 Model comparison (p. 224)
  • 7.6 Summary report (p. 226)
  • 8 Web clickstream analysis (p. 229)
  • 8.1 Objectives of the analysis (p. 229)
  • 8.2 Description of the data (p. 229)
  • 8.3 Exploratory data analysis (p. 232)
  • 8.4 Model building (p. 238)
  • 8.4.1 Sequence rules (p. 238)
  • 8.4.2 Link analysis (p. 242)
  • 8.4.3 Probabilistic expert systems (p. 244)
  • 8.4.4 Markov chains (p. 245)
  • 8.5 Model comparison (p. 250)
  • 8.6 Summary report (p. 252)
  • 9 Profiling website visitors (p. 255)
  • 9.1 Objectives of the analysis (p. 255)
  • 9.2 Description of the data (p. 255)
  • 9.3 Exploratory analysis (p. 258)
  • 9.4 Model building (p. 258)
  • 9.4.1 Cluster analysis (p. 258)
  • 9.4.2 Kohonen maps (p. 262)
  • 9.5 Model comparison (p. 264)
  • 9.6 Summary report (p. 271)
  • 10 Customer relationship management (p. 273)
  • 10.1 Objectives of the analysis (p. 273)
  • 10.2 Description of the data (p. 273)
  • 10.3 Exploratory data analysis (p. 275)
  • 10.4 Model building (p. 278)
  • 10.4.1 Logistic regression models (p. 278)
  • 10.4.2 Radial basis function networks (p. 280)
  • 10.4.3 Classification tree models (p. 281)
  • 10.4.4 Nearest-neighbour models (p. 285)
  • 10.5 Model comparison (p. 286)
  • 10.6 Summary report (p. 290)
  • 11 Credit scoring (p. 293)
  • 11.1 Objectives of the analysis (p. 293)
  • 11.2 Description of the data (p. 294)
  • 11.3 Exploratory data analysis (p. 296)
  • 11.4 Model building (p. 299)
  • 11.4.1 Logistic regression models (p. 299)
  • 11.4.2 Classification tree models (p. 303)
  • 11.4.3 Multilayer perceptron models (p. 314)
  • 11.5 Model comparison (p. 314)
  • 11.6 Summary report (p. 319)
  • 12 Forecasting television audience (p. 323)
  • 12.1 Objectives of the analysis (p. 323)
  • 12.2 Description of the data (p. 324)
  • 12.3 Exploratory data analysis (p. 327)
  • 12.4 Model building (p. 337)
  • 12.5 Model comparison (p. 347)
  • 12.6 Summary report (p. 350)
  • Bibliography (p. 353)
  • Index (p. 357)