Skip to main content

Posts

Showing posts with the label Data visualization

Data preparation: Dealing with duplicates in data(practical example)

In previous post i have said that duplicate values create biased while training model  and have explained how to deal with duplicate values step by step. if you have not read that post make sure you read that (link-- https://www.blogger.com/blog/post/edit/950003373928384423/4720144211787135428 ) as this post going to be practical implementation of theory explained in that post . or just continue reading as i shall explain while showing how it actually work in practical. Steps-- - 1. finding duplicate values through data exploration ---     before we start finding duplicates we need to first import dataset and set columns     so ones data is imported we can start exploring it use dtypes, describe and other methods to explore data   see the image below and observe carefully you will find that we have printed two shapes first one being      shape of  entire data frame and the other one being shape of unique customer_id . and what ...

Data preparation: Treating errors and outliers

Errors are outliers are the those values in dataset which are far away from mean value of that column in which they exist or we can say they are very small or large as compared to other values in that column. suppose we are observing column of price and we see value that is ten times bigger then other values then it could be error or outlier. to identify either it is error or outlier we need to apply some domain knowledge or we can say field knowledge. outliers are mainly caused by variability in measurement or experimental errors. but sometimes they could be useful and help us understand data better if we apply domain knowledge. for example we find a outlier in car price of cars data set then we check its other features and we found  that that car is luxury car then we understand it was outlier but a useful one and it was not just a error. Dealing with errors/outliers--- 1. Detect errors/outliers -- use data exploration to identify errors /outliers 2. methods to identify errors a...

What Frequency Tables for categorical values in dataframe ?

Frequency tables are tables that shows how frequently various categories of categorical variables occur in data and how many different categories are there and which are those categories,  it is also useful for classification to find the frequency of each category of label variable(column) .  this help us to separate helpful categories from not so helpful categories.  suppose some category in categorical variable occurs just ones or twice  then it is not going to be helpful from statistical point of view . Lets see how to make frequency tables--- first i have downloaded auto_prices data set ,then i have taken out come categorical columns and created list of those columns ,this list along with dataset is passed to the count unique function . that function simply loop through each column in the list and  counts  number of times each unique value occurs in  column and finally prints the same. above code gives following frequency table--- Examining classe...

Data visualization for classification

The aim of  Data visualization is some what different for classification as compared to the regression , in classification we have to find how different attributes (numeric and categorical ) are related to  categorical labels or we can say with different categories  of labels. >>following are some techniques used for Data visualization for classification ---- Visualize class separation using numeric feature - ---                                                                             goal of visualizing data for classification is to understand which feature is useful for class separation.                                                   ...