Ask Homework Help/Study Tips Expert

PART 1: CLASSIFICATION

This part of the assignment is concerned with the file:

/KDrive/SEH/SCSIT/Students/Courses/COSC2111/DataMining/ data/other/bank-balanced1.csv.

There is a description of the data in the file bank-names.txt in the same directory. [bank-balanced1.csv is a subset of bank-full.csv]. The main goal is to achieve the highest classification accuracy with the lowest amount of overfitting.

1. Run the following classifiers, with the default parameters, on this data: ZeroR, OneR, J48, IBK and construct a table of the training and cross-validation errors. You can get the training error by selecting "Use training set" as the test option. What do you conclude from these results?

Run No

Classifier

Parameters

Parameters

Training

Error

Cross-valid

Error

Over-

Fitting

1

.

ZeroR

.

None

.

30.0%

.

30.0%

.

None

2. Using the J48 classifier, can you find a combination of the C and M parameter values that minimizes the amount of overfitting? Include the results of your best five runs, including the parameter values, in your table of results.

3. Reset J48 parameters to their default values. What is the effect of lowering the number of examples in the training set? Include your runs in your table of re- sults.

4. Using the IBk classifier, can you find the value of k that minimizes the amount of overfitting? Include your runs in your table of results.

5. Try a number of other classifiers. Aside from ZeroR, which classifiers are best and worst in terms of predictive accuracy? Include 5 runs in your table of results.

6. Compare the accuracy of ZeroR, OneR and J48. What do you conclude?

7. What golden nuggets did you find, if any?

8. [OPTIONAL] Use an attribute selection algorithm to get a reduced attribute set. How does the accuracy on the reduced set compare with the accuracy on the full set?

Report Length: Up to two pages.

PART 2: NUMERIC PREDICTION

Numeric Prediction of the balance attribute in the bank data of part 1. The main goal is to achieve the lowest mean absolute error with the lowest amount of overfitting.

1. Run the following classifers, with default parameters, on this data: ZeroR, MP5, IBk and construct a table of the training and cross-validation errors. You may want to turn on "Output Predictions" to get a better sense of the magnitude of the error on each example. What do you conclude from these results?

2. Explore different parameter settings for M5P and IBk. Which values give the best performance in terms of predictive accuracy and overfitting. Include the results of the best five runs in your table of results.

3. Investigate three other classifiers for numeric prediction and their associated pa- rameters. Include your best five runs in your table of results. Which classifier gives the best performance in terms of predictive accuracy and overfitting?

4. What golden nuggets did you find, if any?

Report Length Up to one page.

PART 3: CLUSTERING

Clustering of the bank data of part 1. For this part use only the attributes age, marital, education, and balance.

The aim is determine the number of clusters in the data and assess whether any of the clusters are meaningful.

1. Run the Kmeans clustering algorithm on this data for the following values of K: 1,2,3,4,5,10,20. Analyse the resulting clusters. What do you conclude?

2. Choose a value of K and run the algorithm with different seeds. What is the effect of changing the seed?

3. Run the EM algorithm on this data with the default parameters and describe the output.

4. The EM algorithm can be quite sensitive to whether the data is normalized or not. Use the weka normalize filter (Preprocess --> Filter --> unsupervised --> normalize) to normalize the numeric attributes. What difference does this make to the clus- tering runs?

5. The algorithm can be quite sensitive to the values of minLogLikelihoodImprove- mentCV minStdDev and minLogLikelihoodImprovementIterating, Explore the effect of changing these values. What do you conclude?

6. How many clusters do you think are in the data? Give an English language description of one of them.

7. Compare the use of Kmeans and EM for these clustering tasks. Which do you think is best? Why?

8. What golden nuggets did you find, if any?
Report Length Up to one page.

PART 4: ASSOCIATION FINDING

Association finding in the files supermarket1.arff and supermarket2.arff in the folder
/KDrive/SEH/SCSIT/Students/Courses/COSC2111/DataMining/data/arff.

The main aim is to determine whether there are any significant associations in the data.

These files contain the same details of shopping transactions represented in two different ways. You can use a text viewer to look at the files.

1. What is the difference in representations?

2. Load the file supermarket1.arff into weka and run the Apriori algorithm on this data. You might need to restrict the number of attributes and/or the number of examples. What significant associations can you find?

3. Explore different possibilities of the metric type and associated parameters. What do you find?

4. Load the file supermarket22.arff into weka and run the Apriori algorithm on this data. What do you find?

5. Explore different possibilities of the metric type and associated parameters. What do you find?

6. Try the other associators. What are the differences to Apriori?

7. What golden nuggets did you find, if any?

8. [OPTIONAL] Can you find any meaningful associations in the bank data?

Report Length Up to one page.

Homework Help/Study Tips, Others

  • Category:- Homework Help/Study Tips
  • Reference No.:- M93066840
  • Price:- $50

Priced at Now at $50, Verified Solution

Have any Question?


Related Questions in Homework Help/Study Tips

Review the website airmail service from the smithsonian

Review the website Airmail Service from the Smithsonian National Postal Museum that is dedicated to the history of the U.S. Air Mail Service. Go to the Airmail in America link and explore the additional tabs along the le ...

Read the article frank whittle and the race for the jet

Read the article Frank Whittle and the Race for the Jet from "Historynet" describing the historical influences of Sir Frank Whittle and his early work contributions to jet engine technologies. Prepare a presentation high ...

Overviewnow that we have had an introduction to the context

Overview Now that we have had an introduction to the context of Jesus' life and an overview of the Biblical gospels, we are now ready to take a look at the earliest gospel written about Jesus - the Gospel of Mark. In thi ...

Fitness projectstudents will design and implement a six

Fitness Project Students will design and implement a six week long fitness program for a family member, friend or co-worker. The fitness program will be based on concepts discussed in class. Students will provide justifi ...

Read grand canyon collision - the greatest commercial air

Read Grand Canyon Collision - The greatest commercial air tragedy of its day! from doney, which details the circumstances surrounding one of the most prolific aircraft accidents of all time-the June 1956 mid-air collisio ...

Qestion anti-trustprior to completing the assignment

Question: Anti-Trust Prior to completing the assignment, review Chapter 4 of your course text. You are a manager with 5 years of experience and need to write a report for senior management on how your firm can avoid the ...

Question how has the patient and affordable care act of

Question: How has the Patient and Affordable Care Act of 2010 (the "Health Care Reform Act") reshaped financial arrangements between hospitals, physicians, and other providers with Medicare making a single payment for al ...

Plate tectonicsthe learning objectives for chapter 2 and

Plate Tectonics The Learning Objectives for Chapter 2 and this web quest is to learn about and become familiar with: Plate Boundary Types Plate Boundary Interactions Plate Tectonic Map of the World Past Plate Movement an ...

Question critical case for billing amp codingcomplete the

Question: Critical Case for Billing & Coding Complete the Critical Case for Billing & Coding simulation within the LearnScape platform. You will need to create a single Microsoft Word file and save it to your computer. A ...

Review the cba provided in the resources section between

Review the CBA provided in the resources section between the Trustees of Columbia University and Local 2110 International Union of Technical, Office, and Professional Workers. Describe how this is similar to a "contract" ...

  • 4,153,160 Questions Asked
  • 13,132 Experts
  • 2,558,936 Questions Answered

Ask Experts for help!!

Looking for Assignment Help?

Start excelling in your Courses, Get help with Assignment

Write us your full requirement for evaluation and you will receive response within 20 minutes turnaround time.

Ask Now Help with Problems, Get a Best Answer

Why might a bank avoid the use of interest rate swaps even

Why might a bank avoid the use of interest rate swaps, even when the institution is exposed to significant interest rate

Describe the difference between zero coupon bonds and

Describe the difference between zero coupon bonds and coupon bonds. Under what conditions will a coupon bond sell at a p

Compute the present value of an annuity of 880 per year

Compute the present value of an annuity of $ 880 per year for 16 years, given a discount rate of 6 percent per annum. As

Compute the present value of an 1150 payment made in ten

Compute the present value of an $1,150 payment made in ten years when the discount rate is 12 percent. (Do not round int

Compute the present value of an annuity of 699 per year

Compute the present value of an annuity of $ 699 per year for 19 years, given a discount rate of 6 percent per annum. As