Question Bank- Data Pre-processing and Exploration

  1. Suppose we have a dataset of employees’ working experience (in years), and one of the attributes in the dataset is “exp_years”, representing the number of years of professional experience each employee has. The values of “exp_years” in the dataset are as follows (in increasing order) : 1, 2, 3, 5, 6, 7, 9, 11, 14, 15, 18, 21, 23, 25, 28.
    1. Calculate the Z-score for the experience value 15 using the given dataset and standardize it.
    2. Calculate the mean of “exp_years”
    3. Calculate the standard deviation of “exp_years”
    4. Calculate the median and mode of the dataset
    5. Draw a box plot for the “exp_years” values

2. Sketch concept hierarchy for location and product category

3.Suppose a group of daily humidity percentages for a location has been sorted in increasing order as follows: 35, 37, 40, 42, 45, 47, 50, 52, 55, 58, 60, 62. We want to partition these humidity records into four bins using the equal-frequency partitioning method. Then, we’ll perform data smoothing using bin mean, bin median, and bin boundary methods for each bin.

a. Perform equal-frequency partitioning to divide the given humidity records into four bins.

b. Apply data smoothing using: Bin Mean, Bin Median, Bin Boundary for each of the four bins. c. Present the final results with smoothed values for each bin using all three methods.

4.Scenario: A company wants to analyze the monthly sales revenue (in thousands of dollars) from its four different products over the year.

Given Sales Data (in thousands of dollars):

5, 8, 12, 15, 17, 21, 25, 28, 30, 34, 37, 40

Task:

Perform equal-frequency partitioning to divide the given sales data into four bins.

Apply data smoothing using bin mean, median, and boundary methods for each bin. Present the final results with smoothed values for each bin.

5. Sketch concept hierarchy for time and product category.

6. Calculate the range and mid-range of the following values: 12, 18, 24, 29, 35, 41, 47, 53, 59, 64, 71, 78, 83, 89, 95

7. Suppose a group of sales price records has been sorted as follows: 6, 9, 12, 13, 15, 25, 50, 70, 72, 94, 204, 232. Partition them into 3 bins by equal -frequency partitioning method and apply smoothing by bin mean, median and boundary.

8. Suppose a group of temperature records (in degrees Celsius) for a city has been sorted as follows: 2, 4, 7, 8, 10, 14, 16, 20, 22, 25, 28, 30. We want to partition these temperature records into four bins by equal-frequency partitioning method. Then, we’ll perform data smoothing using bin mean, median, and boundary methods for each bin.

Your task is to:

  1. Perform equal-frequency partitioning to divide the given temperature records into four bins.
  2. Apply data smoothing using bin mean, median, and boundary methods for each bin.
  3. Present the final results with smoothed values for each bin.

Please show your step-by-step calculations and provide the final answers.

9. Suppose we have a dataset of students’ exam scores, and one of the attributes in the dataset is “test_age,” representing the age of the students when they took the test. The ages in the dataset are as follows (in increasing order): 13, 15, 16, 16, 19, 20, 23, 29, 33, 41, 44, 53, 62, 69, 72.

perform Z-score normalization on the “test_age” attribute to standardize the values and bring them onto a common scale.

Your task is to calculate the Z-score for the age value 45 using the given dataset and answer the following questions:

  1. Calculate mean age of the students in the dataset
  2. Calculate standard deviation of the age values in the dataset
  3. Calculate Z-score of the age value 45 after performing Z-score normalization

Please show your step-by-step calculations and provide the final answers.

10. Suppose that data for analysis includes the attribute age of a survey conducted on social media use. The age values for data tuples are (in increasing order):

13,15,16,16,19,20,20,21,22,22,25,25,25,25,30,33,33,35,35,35,35,36,40,45,46,52,70

Draw box plot of the data Calculate mean of the data. Calculate median of the data . Calculate mode of  the data. Comment on modality. Give the five- point summary of the data.