Welcome to the final project repository for NLP course! Dive into our intricate analysis of word pairs using a plethora of datasets. Uncover contextual nuances and compare how computational models stand up against human judgment in the realm of semantic similarity.
This project aims to explore semantic relationships between various word pairs using datasets like MC, Senseval-2, RG, and WSD353. It uses various methods and packages such as WikiSim, FastText ect. Through detailed visualizations and metrics, we analyze the effectiveness of FastText embeddings and other NLP techniques in capturing the essence of human judgments.
- Python 3.x
- pip
- WikiSim (https://github.com/asajadi/wikisim)
Clone the repository:
git clone https://github.com/Mobusshar/NLP_Final_Project.git
cd NLP_Final_ProjectThis code will engage the datasets, run the analysis, and generate insightful visualizations.
Initialize WikiSim to obtain required similarities.
We employ the following datasets for our analysis:
- MC Dataset: Delve into word pairs of varied semantic relationships.
- RG Dataset: Explore an extensive range of word pair spectrums.
- WSD353 Dataset: Engage in a deep dive across an extensive semantic landscape.
- Senseval-2: Use Senseval to test Wordsense Disambiguation methods.
Navigate to the results directory to view generated visualizations and insights. Witness the side-by-side comparisons of model-generated similarity scores by human judgments.
Want to contribute? Please follow the contributing guidelines.
The analysis is backed by extensive literature. Kindly refer to the references section for a detailed list of academic and technical sources that guided this project.
MIT License. For further details, refer to the LICENSE file.