Skip to content

Repository files navigation

Visual Question Answering Visualizer (VQA-VIZ)

An easy-to-use app to visualise attentions of various VQA models.

top 7 predictions

• Models
• Requirements
• Installation
• How to run
• How to use
• Contributing
• Acknowledgements

Models

• MFB - Multi-modal Factorized Bilinear Pooling with Co-Attention Learning for Visual Question Answering
Zhou Yu, Jun Yu, Jianping Fan, Dacheng Tao
Arxiv

• (Coming soon) MCAN - Deep Modular Co-Attention Networks for Visual Question Answering
Zhou Yu, Jun Yu, Yuhao Cui, Dacheng Tao, Qi Tian
Arvix

Requirements

Please check the requirements.txt file for the version numbers.

  1. torchvision
  2. seaborn
  3. pandas
  4. matplotlib
  5. dotmap
  6. streamlit
  7. numpy
  8. torch
  9. torchvision
  10. Pillow
  11. PyYAML
  12. opencv_python

Installation

  1. Install Anaconda
  2. Clone this repository and cd into it.
    git clone https://github.com/apugoneappu/vqa_visualise.git && cd vqa_visualise
  3. In a new environment (new_env)
    pip install -r requirements.txt

How to run

From the directory of this repository, do the following -

  1. conda activate new_env
  2. streamlit run vqa_input.py
  3. In a browser tab, open the Network URL displayed in your terminal.

Done! 🎉

How to use

input page top 7 predictions image attentions text attentions

Contributing

First of all, thank you for wanting to contribute to this work! I will try and make your job as easy as possible.

Acknowledgements

This repository has been built by modifying the OpenVQA repository.

I would also like to thank Yash Khandelwal, Nikhil Shah and Chinmay Singh for their support and amazing suggestions!

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages