Repository navigation
Estimating document topic distributions for short social media texts #2494
Unanswered
oromiaGodanna
asked this question in
Q&A
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Hi, I’m using BERTopic on millions of short social-media posts and would appreciate advice on interpreting/handling the probability distribution outputs.
Setup:
all-MiniLM-L12-v2, normalizedcalculate_probabilities=TrueCountVectorizer(stop_words="english", ngram_range=(1,2), min_df=3, max_df=0.8)My goal is to get a reasonable document-level probability distribution across the discovered clusters/topics.
The issue:
-1by HDBSCAN (about 60 %); I do outlier reduction using probabilitiescalculate_probabilities=True, many documents have weak probability mass spread across a very large number of clusters (non-zero probability across more than 50% of the clusters)approximate_distribution, the distributions can also be quite diffuse, with many non-zero topic probabilities per document but much better than HDBSCAN. Then when I compare the primary topic assigned by this approximate distribution with HDBSCAN-assigned topics, I get only about 35% overlap. Additionally, usinguse_embedding_model=Truehere doesn't improve much.My question:
For short social-media posts, what is the recommended way to obtain a usable probability distribution per document in BERTopic? (Usable means that the probability distribution per document is spread within a reasonable number of clusters or topics, given that the mean number of tokens per document is about 10).
Any guidance on best practices or diagnostics for this case would be appreciated.
All reactions