
2000 Olympics diving Difficulty/Average Score distribution plot
Background
This week we received a dataset of 2000 Olympics diving dataset. When I first got this dataset I have no clue what I want to make for visualization, but the week before I saw my classmate did a visualization with a network plot so I decided to do something similar.
Data
The dataset includes all the dives in the 2000 Olympics from preliminary to final. It has the score, who was the judge, what country was that judge from, the average score, difficulty, and the diver's rank.
R
After having my dataset all prepared, I started the coding part. My thought process was to showcase the distribution by difficulty to the average score. The first thing I did was to give them a category. So I calculated the quantile of difficulty and score and then put them into five groups. Next, I calculated how many observations in each category. Then it's the most important thing, I created two nodes for two distributions. Finally, I used htmlwidgets to create the distribution plot.
Summary
From the result, we can see that most of the divers choose the difficulty between 2.6-3.0 with a total of 4935 dives. Only 42 dives are 3.5 or above. Also, we can see from the score distribution, it matches pretty well to the quantile. For me, this is a good experience for me to learn a new way to create a visualization.