Hi everyone, author from the SGG-Benchmark codebase here.
First thank you for your repository @ChocoWu, very valuable work for the community! And thank you for highlighting my work :)
I have a few insight on how to provide a guidebook to make it easier for new people to jump into the field of SGG. In my opinion, there is a few bottlenecks that makes it very difficult for people to build on top of current work and/or use SGG approaches in real world applications:
-
The lack of annotations standard and pipeline. SGG is mostly a supervised task and to use most models you need to annotate relations, which may not be straightforward. I am providing an annotation tool (see https://github.com/Maelic/SGG-Annotate) and I have been working on implementing a new annotation standard for the community directly into the coco API (see https://github.com/Maelic/cocoapi). But to make it easier, providing a semi-automated pipeline using for instance SAM3 and VLMs to quickly annotate images could be nice. Also, providing insights on how to create a new dataset and choose relation categories would be valuable (i.e. not all predicates are exclusive etc). Finally, implementing a fair evaluation strategy (see https://arxiv.org/abs/2407.09216), and good metrics (see https://arxiv.org/abs/2404.09616) directly into the coco API, for both panoptic and box SGG methods could be nice. Having a discussion on setting a standard for the metrics will also be very beneficial, in recent years a lot of people proposed better metrics than classical Recall@K and meanRecall@K, however new papers are still using these metrics.
-
Regarding codebases, there has been some effort in the past to standardize the code and implement baseline approaches in a single codebase (huge shoutout to Kaihua, the field of SGG wouldn't be the same without you 🙏). However, I feel that now the field has become divided again (for instance between One-Stage and Two-Stage approaches, panoptic approaches or VLM-based approaches etc) whereas, at the end of the day, most of the approaches share a common ground (either the underlying detector or the feature extraction etc). As a result, I think it would be possible to integrate some SGG approaches directly on top of up-to-date codebase such as https://github.com/ultralytics/ultralytics. The ultralytics codebase is simple, easy-to-use and provides a lot of guidebook for beginners. In addition, ultralytics incorporate object detection models (for 2D SGG) AND segmentation models (for panoptic SGG) as well as open-vocabulary/VLM models (for Open-Vocabulary SGG). I have myself started to implement SGG metrics and evaluation code into the ultralytics codebase (see https://github.com/Maelic/YOLO-Rel), if you are interested to help please contact me!
-
The lack of deployment support. What makes me particularly sad today is that, after thousands of papers and really cool work in SGG, nobody is using SGG models in real-world applications 😞. I think the main root of the problem (in addition to previous points) is that we are lacking checkpoints in .onnx or even in .safetensors forms that could be used with the ONNX API or HF for easy deployment. There is also no code (to my knowledge) that provides a guidebook on how to export a SGG model from plain pytorch to ONNX, TensorRT etc... which are the main formats used in production. I know that as scientists our work should not be on the deployment but as a community, if nobody is providing such tools, then nobody will use SGG models in their applications and then the interest for the field will slowly die. Codebases such as https://github.com/ultralytics/ultralytics are providing a lot of deployment tools, which could be a starting point for us, but if you have a better idea to help people integrate SGG approaches in their applications please feel free to share!
That's all for me for today, I hope my work and insights will be able to spark really cool discussion in the SGG community 🚀
Maelic
p.s. I won't be able to make it for your WACV workshop @ChocoWu, but please say hello to Azade from me :)
Hi everyone, author from the SGG-Benchmark codebase here.
First thank you for your repository @ChocoWu, very valuable work for the community! And thank you for highlighting my work :)
I have a few insight on how to provide a guidebook to make it easier for new people to jump into the field of SGG. In my opinion, there is a few bottlenecks that makes it very difficult for people to build on top of current work and/or use SGG approaches in real world applications:
The lack of annotations standard and pipeline. SGG is mostly a supervised task and to use most models you need to annotate relations, which may not be straightforward. I am providing an annotation tool (see https://github.com/Maelic/SGG-Annotate) and I have been working on implementing a new annotation standard for the community directly into the coco API (see https://github.com/Maelic/cocoapi). But to make it easier, providing a semi-automated pipeline using for instance SAM3 and VLMs to quickly annotate images could be nice. Also, providing insights on how to create a new dataset and choose relation categories would be valuable (i.e. not all predicates are exclusive etc). Finally, implementing a fair evaluation strategy (see https://arxiv.org/abs/2407.09216), and good metrics (see https://arxiv.org/abs/2404.09616) directly into the coco API, for both panoptic and box SGG methods could be nice. Having a discussion on setting a standard for the metrics will also be very beneficial, in recent years a lot of people proposed better metrics than classical Recall@K and meanRecall@K, however new papers are still using these metrics.
Regarding codebases, there has been some effort in the past to standardize the code and implement baseline approaches in a single codebase (huge shoutout to Kaihua, the field of SGG wouldn't be the same without you 🙏). However, I feel that now the field has become divided again (for instance between One-Stage and Two-Stage approaches, panoptic approaches or VLM-based approaches etc) whereas, at the end of the day, most of the approaches share a common ground (either the underlying detector or the feature extraction etc). As a result, I think it would be possible to integrate some SGG approaches directly on top of up-to-date codebase such as https://github.com/ultralytics/ultralytics. The ultralytics codebase is simple, easy-to-use and provides a lot of guidebook for beginners. In addition, ultralytics incorporate object detection models (for 2D SGG) AND segmentation models (for panoptic SGG) as well as open-vocabulary/VLM models (for Open-Vocabulary SGG). I have myself started to implement SGG metrics and evaluation code into the ultralytics codebase (see https://github.com/Maelic/YOLO-Rel), if you are interested to help please contact me!
The lack of deployment support. What makes me particularly sad today is that, after thousands of papers and really cool work in SGG, nobody is using SGG models in real-world applications 😞. I think the main root of the problem (in addition to previous points) is that we are lacking checkpoints in .onnx or even in .safetensors forms that could be used with the ONNX API or HF for easy deployment. There is also no code (to my knowledge) that provides a guidebook on how to export a SGG model from plain pytorch to ONNX, TensorRT etc... which are the main formats used in production. I know that as scientists our work should not be on the deployment but as a community, if nobody is providing such tools, then nobody will use SGG models in their applications and then the interest for the field will slowly die. Codebases such as https://github.com/ultralytics/ultralytics are providing a lot of deployment tools, which could be a starting point for us, but if you have a better idea to help people integrate SGG approaches in their applications please feel free to share!
That's all for me for today, I hope my work and insights will be able to spark really cool discussion in the SGG community 🚀
Maelic
p.s. I won't be able to make it for your WACV workshop @ChocoWu, but please say hello to Azade from me :)