|
Hi Maarten, I hope you're well! This may be a silly question but I've been having a little bit of trouble understanding exactly what's going on under the hood with BERTopic regarding choosing cluster labels, and wondered if someone can clarify. Please let me know if this has been asked before. I've found a counter-example to an assumption I was making (namely, if I embed my documents, reduce them, and cluster them with HDBSCAN, then the cluster labels I get from HDBSCAN will correspond to the topic labels I later get from running the full BERTopic algorithm with the same embedding, dimensionality reduction, and clustering models), by choosing a particular set of parameters on the documents I'm analysing: import numpy as np
import pandas as pd
from bertopic import BERTopic
from hdbscan import HDBSCAN
from sentence_transformers import SentenceTransformer
from sklearn.decomposition import PCA
docs = pd.read_csv("docs.csv")["0"].tolist()
# Embed
embedding_model = SentenceTransformer("all-MiniLM-L6-v2")
embeddings = embedding_model.encode(docs)
# Reduce
reducer = PCA(n_components=13, random_state=42)
reduced_embeddings = reducer.fit_transform(embeddings)
# Cluster
clusterer = HDBSCAN(
min_cluster_size=13,
min_samples=8,
metric="euclidean",
cluster_selection_method="eom",
prediction_data=True,
)
clustered = clusterer.fit(reduced_embeddings)
# Show unique cluster labels
np.unique(clustered.labels_)
As you can see, all my documents have been assigned the label Now we plug the same combination of models into BERTopic: topic_model = BERTopic(
embedding_model=embedding_model,
umap_model=reducer,
hdbscan_model=clusterer,
# nr_topics="auto",
)
topics, probs = topic_model.fit_transform(docs, embeddings)
topic_info = topic_model.get_topic_info()
topic_info[["Topic", "Count"]]
Now, despite choosing the same embedding, dimensionality reduction, and clustering models to pass to I hope this makes sense. Thank you, |
Replies: 3 comments 3 replies
|
iirc, it should be indeed the same output as long as both the reducer and clusterer will generate the same output every time. Do you get the same output each time when you run them without BERTopic (so your first example)?? |
|
Hopefully this is a slightly more reproducible example (if a bit contrived) using a public dataset: import numpy as np
import pandas as pd
from bertopic import BERTopic
from hdbscan import HDBSCAN
from sentence_transformers import SentenceTransformer
from sklearn.datasets import fetch_20newsgroups
from sklearn.decomposition import PCA
docs = fetch_20newsgroups(
subset='all', remove=('headers', 'footers', 'quotes')
)['data'][:2000]# Embed
embedding_model = SentenceTransformer("all-MiniLM-L6-v2")
embeddings = embedding_model.encode(docs)
# Reduce
reducer = PCA(n_components=13, random_state=42)
reduced_embeddings = reducer.fit_transform(embeddings)
# Cluster
clusterer = HDBSCAN(
min_cluster_size=60,
min_samples=30,
metric="euclidean",
cluster_selection_method="eom",
prediction_data=True,
)
clustered = clusterer.fit(reduced_embeddings)
# Show unique cluster labels
pd.DataFrame(
np.unique(clustered.labels_, return_counts=True)
).T.rename(columns={0: "Topic", 1: "Count"})
topic_model = BERTopic(
embedding_model=embedding_model,
umap_model=reducer,
hdbscan_model=clusterer,
calculate_probabilities=True,
# nr_topics="auto",
)
topics, probs = topic_model.fit_transform(docs, embeddings)
topic_info = topic_model.get_topic_info()
topic_info[["Topic", "Count"]]
class EmptyReducer:
def __init__(self):
pass
def fit(self, embeddings):
return embeddings
def fit_transform(self, embeddings):
return embeddings
def transform(self, embeddings):
return embeddings
topic_model = BERTopic(
embedding_model=embedding_model,
umap_model=EmptyReducer(),
hdbscan_model=clusterer,
calculate_probabilities=True,
# nr_topics="auto",
)
topics, probs = topic_model.fit_transform(docs, reduced_embeddings)
topic_info = topic_model.get_topic_info()
topic_info[["Topic", "Count"]]
All three of these outputs are slightly different (the only difference between the first and third is that clusters 0 and 1 get swapped around I suppose?). But the second one has different numbers of documents assigned to each cluster. This is all using |
|
Hi @MaartenGr , I found the difference between what If instead of: reducer = PCA(n_components=13, random_state=42)
reduced_embeddings = reducer.fit_transform(embeddings)I do: reducer = PCA(n_components=13, random_state=42)
reducer.fit(embeddings)
reduced_embeddings = reducer.transform(embeddings)Then the outputs of both approaches will align, as this latter one matches the behaviour of |
Hi @MaartenGr , I found the difference between what
BERTopic()does and what I was doing.If instead of:
I do:
Then the outputs of both approaches will align, as this latter one matches the behaviour of
BERTopic._reduce_dimensionality(). Was expectingself.umap_model.fit(embeddings)followed byumap_embeddings = self.umap_model.transform(embeddings)to be equivalent toself.umap_model.fit_transform(embeddings), but apparently this is not the case!