Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
56 changes: 56 additions & 0 deletions project/.ipynb_checkpoints/report-template-checkpoint.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,56 @@
# Report: Predict Bike Sharing Demand with AutoGluon Solution
#### AJAYI JOHN

## Initial Training
### What did you realize when you tried to submit your predictions? What changes were needed to the output of the predictor to submit your results?
TODO: Add your explanation
Kaggle rejects negative score values, hence, the negative values was set to zero before submission

### What was the top ranked model that performed?
TODO: Add your explanation
The top ranked model that performed is the "WeightedEnsemble_L3" with a score of and RMSE value of -53.055. RMSE evaluation metric was used. Hence, the lower the error, the better the model.

## Exploratory data analysis and feature creation
### What did the exploratory analysis find and how did you add additional features?
TODO: Add your explanation
The temp and atemp features are almost perfectly skewed, indicating that the temperature range is almost the same all through the week.
The causal users, windspeed, registered and count columns are left-skewed. This is an indicator that do not have particular periods of the day to lend a bike.
The other columns are simply discrete values

### How much better did your model preform after adding additional features and why do you think that is?
TODO: Add your explanation
After adding additional features, the model inproved significantly. This included also converting certain features to categorical variables.
From an initial RMSE value of -53 to RMSE value of -30.1837. Also, the kaggle score improved from 1.79030 to 0.67749

## Hyper parameter tuning
### How much better did your model preform after trying different hyper parameters?
TODO: Add your explanation
Tuning of some of the hyerparameters also improved the model's performace
Some of the hyperparameters that were tuned included the 'number of boost rounds', 'learning rate', 'num_leaves', 'max_depth'
The hyperparameter tuning moved the model from a previous RMSE value of -30 to -32.78335 though with a significant increase in the kaggle score wwhich went from 0.67749 to 0.44964. This is probably an indication of the required metric needed for evaluating the model.

### If you were given more time with this dataset, where do you think you would spend more time?
TODO: Add your explanation
Given more time with this dataset, I would focus on how working more on the feature engineering and also on the hyperparameter tuning. These two tends to have a significant effect on the model's performance.

### Create a table with the models you ran, the hyperparameters modified, and the kaggle score.
|model|hpo1|hpo2|hpo3|score|
|--|--|--|--|--|
|initial|?|?|?|?|
|add_features|?|?|?|?|
|hpo|?|?|?|?|

### Create a line plot showing the top model score for the three (or more) training runs during the project.

TODO: Replace the image below with your own.

![model_train_score.png](img/model_train_score.png)

### Create a line plot showing the top kaggle score for the three (or more) prediction submissions during the project.

TODO: Replace the image below with your own.

![model_test_score.png](img/model_test_score.png)

## Summary
TODO: Add your explanation
Binary file modified project/img/model_test_score.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file modified project/img/model_train_score.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added project/model_test_score.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
4,015 changes: 3,868 additions & 147 deletions project/project-template.ipynb

Large diffs are not rendered by default.

33 changes: 19 additions & 14 deletions project/report-template.md
Original file line number Diff line number Diff line change
@@ -1,39 +1,44 @@
# Report: Predict Bike Sharing Demand with AutoGluon Solution
#### NAME HERE
#### AJAYI JOHN

## Initial Training
### What did you realize when you tried to submit your predictions? What changes were needed to the output of the predictor to submit your results?
TODO: Add your explanation
Kaggle rejects negative score values, hence, the negative values was set to zero before submission

### What was the top ranked model that performed?
TODO: Add your explanation
The top ranked model that performed is the "WeightedEnsemble_L3" with a score of and RMSE value of -53.055. RMSE evaluation metric was used. Hence, the lower the error, the better the model.

## Exploratory data analysis and feature creation
### What did the exploratory analysis find and how did you add additional features?
TODO: Add your explanation
The temp and atemp features are almost perfectly skewed, indicating that the temperature range is almost the same all through the week.
The causal users, windspeed, registered and count columns are left-skewed. This is an indicator that do not have particular periods of the day to lend a bike.
The other columns are simply discrete values

### How much better did your model preform after adding additional features and why do you think that is?
TODO: Add your explanation
After adding additional features, the model inproved significantly. This included also converting certain features to categorical variables.
From an initial RMSE value of -53 to RMSE value of -30.1837. Also, the kaggle score improved from 1.79030 to 0.67749

## Hyper parameter tuning
### How much better did your model preform after trying different hyper parameters?
TODO: Add your explanation
Tuning of some of the hyerparameters also improved the model's performace
Some of the hyperparameters that were tuned included the 'number of boost rounds', 'learning rate', 'num_leaves', 'max_depth'
The hyperparameter tuning moved the model from a previous RMSE value of -30 to -32.78335 though with a significant increase in the kaggle score wwhich went from 0.67749 to 0.44964. This is probably an indication of the required metric needed for evaluating the model.

### If you were given more time with this dataset, where do you think you would spend more time?
TODO: Add your explanation
Given more time with this dataset, I would focus on how working more on the feature engineering and also on the hyperparameter tuning. These two tends to have a significant effect on the model's performance.

### Create a table with the models you ran, the hyperparameters modified, and the kaggle score.
|model|hpo1|hpo2|hpo3|score|
|--|--|--|--|--|
|initial|?|?|?|?|
|add_features|?|?|?|?|
|hpo|?|?|?|?|
|model|CAT|XGB|GBM|RF|score|
|--|--|--|--|--|--|
|initial|default|default|default|default|1.79030|
|add_features|default|default|default|default|0.67749|
|hpo|'depth':8, 'l2_leaf_reg':10'|'objective':'reg:pseudohubererror', 'eval_metric':'rmse', 'max_depth':10, 'eta':0.03|'objective':'regression_l1', 'num_boost_round':500, 'num_leaves':50, 'eta':0.001, 'random_state'=32|'n_estimators':500, 'max_depth':8, 'min_samples_split':4, 'min_samples_leaf':3, 'max_features':'auto', 'random_state':32|0.44964|

### Create a line plot showing the top model score for the three (or more) training runs during the project.

TODO: Replace the image below with your own.

![model_train_score.png](img/model_train_score.png)
![model_train_score.png](img/model_train_score.png)default

### Create a line plot showing the top kaggle score for the three (or more) prediction submissions during the project.

Expand All @@ -42,4 +47,4 @@ TODO: Replace the image below with your own.
![model_test_score.png](img/model_test_score.png)

## Summary
TODO: Add your explanation
In conclussion, I was able to create a model that can be used to predict the number of persons that makes use of the bike-sharing service in each day of the week. This will help the organization make useful decision and planning, hence, aiding the smooth running of the business.
Loading