An easy-to-use app to visualise attentions of various VQA models.
• Models
• Requirements
• Installation
• How to run
• How to use
• Contributing
• Acknowledgements
• MFB - Multi-modal Factorized Bilinear Pooling with Co-Attention Learning for Visual Question Answering
Zhou Yu, Jun Yu, Jianping Fan, Dacheng Tao
Arxiv
• (Coming soon) MCAN - Deep Modular Co-Attention Networks for Visual Question Answering
Zhou Yu, Jun Yu, Yuhao Cui, Dacheng Tao, Qi Tian
Arvix
Please check the requirements.txt file for the version numbers.
- torchvision
- seaborn
- pandas
- matplotlib
- dotmap
- streamlit
- numpy
- torch
- torchvision
- Pillow
- PyYAML
- opencv_python
- Install Anaconda
- Clone this repository and cd into it.
git clone https://github.com/apugoneappu/vqa_visualise.git && cd vqa_visualise - In a new environment (
new_env)
pip install -r requirements.txt
From the directory of this repository, do the following -
conda activate new_envstreamlit run vqa_input.py- In a browser tab, open the Network URL displayed in your terminal.
Done! 🎉
First of all, thank you for wanting to contribute to this work! I will try and make your job as easy as possible.
This repository has been built by modifying the OpenVQA repository.
I would also like to thank Yash Khandelwal, Nikhil Shah and Chinmay Singh for their support and amazing suggestions!




