Please use this identifier to cite or link to this item: http://hdl.handle.net/123456789/5879
Title: Data Cleaning in Knowledge Discovery Database (KDD)-Data Mining
Authors: Fauziah Abdul Rahman
Mohammad Ishak Desa
Antoni Wibowo
Norhaidah Abu Haris
Keywords: Data Cleaning
Data Mining
Missing Value
DC Process
Issue Date: 10-Apr-2014
Abstract: Data quality is a main issue in quality information management. Data quality problems occur anywhere in information systems. These problems are solved by data cleaning. Data cleaning (DC)is a process used to determine inaccurate, incomplete or unreasonable data and then improve the quality through correcting of detected errors and omissions. Generally data cleaning reduces errors and improves the data quality. It is well known that the process of correcting errors in data and eliminating bad records are time consuming and involve a tedious process but it cannot be ignored. Various process of DC have been discussed in the previous studies, but there’s no standard or formalized the DC process. Knowledge Discovery Database (KDD) is a tool that enables one to intelligently analyze and explore extensive data for effective decision making. The Cross-Industry Standard Process for Data Mining (CRISP-DM) is one of the KDD methodology often used for this purpose. This paper review and emphasize the important of DC in data preparation. The wrong analysis will probably turn out to be expensive failures. The future works was also being highlighted.
Description: Conference venue: Kuala Lumpur
URI: http://localhost/xmlui/handle/123456789/5879
Appears in Collections:Conference Paper

Files in This Item:
File Description SizeFormat 
Data Cleaning in Knowledge Discovery Database (KDD)-Data Mining.pdf603.77 kBAdobe PDFView/Open    Request a copy


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.